Systems and methods for translocating cargo nucleotide sequences

JP2025510006A5Pending Publication Date: 2026-02-19METAGENOMI INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024549596
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-05
Filing Date
2023-02-23
Publication Date
2026-02-19

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure provides systems and methods for translocating a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid. These systems and methods may include a double-stranded nucleic acid comprising a cargo nucleotide sequence, where the cargo nucleotide sequence is configured to interact with a recombinase complex, an effector complex comprising at least one engineered guide polynucleotide configured to hybridize to the effector and target nucleic acid, and a recombinase complex, where the recombinase complex is configured to recruit the cargo nucleotide to the target nucleic acid site.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] cross reference This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 313,156, filed February 23, 2022, and U.S. Provisional Patent Application No. 63 / 478,689, filed January 5, 2023, each of which is incorporated by reference in its entirety.

[0002] Sequence Listing The contents of the electronic sequence listing (MTG-008WO_SL.xml, size: 223,433 bytes, and creation date: February 23, 2023) are incorporated herein by reference in their entirety. [Background technology]

[0003] Cas enzymes, along with their associated clustered regularly interspaced short palindromic repeats (CRISPR) guide ribonucleic acid (RNA), appear to be widespread (~45% in bacteria, ~84% in archaea) components of the prokaryotic immune system, helping to protect such microorganisms from non-self nucleic acids, such as infectious viruses and plasmids, by CRISPR-RNA-guided nucleic acid cleavage. While deoxyribonucleic acid (DNA) elements encoding CRISPR RNA elements may be relatively conserved in structure and length, their CRISPR-associated (Cas) proteins are highly diverse and contain a wide variety of nucleic acid-interacting domains. Although CRISPR DNA elements have been observed as early as 1987, the programmable endonuclease cleavage capabilities of CRISPR / Cas complexes have only been recognized relatively recently, leading to the use of recombinant CRISPR / Cas systems in a variety of DNA engineering and gene editing applications. Summary of the Invention

[0004] In some aspects, the disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site, the system comprising: a first double-stranded nucleic acid comprising a cargo nucleotide sequence configured to interact with a Tn7-type transposase complex; a Cas effector complex comprising a class 2 V-type Cas effector and an engineered guide polynucleotide configured to hybridize to the target nucleic acid site; and a Tn7-type transposase complex configured to bind to the Cas effector complex, the Tn7-type transposase complex comprising a TnsB subunit. In some embodiments, the cargo nucleotide sequence is flanked by a left transposase recognition sequence and a right transposase recognition sequence. In some embodiments, the system further comprises a second double-stranded nucleic acid comprising the target nucleic acid site. In some embodiments, the system further comprises a PAM sequence compatible with the nuclease flanking the target nucleic acid site. In some embodiments, the PAM sequence is located 3' of the target nucleic acid site. In some embodiments, the PAM sequence is located 5' of the target nucleic acid site. In some embodiments, the engineered guide polynucleotide is configured to bind to the class 2 V-type Cas effector. In some embodiments, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least 80% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220, or a variant thereof. In some embodiments, the TnsB subunit comprises a polypeptide comprising a sequence having at least 80% identity to SEQ ID NOs: 2, 13, 17, and 65, or a variant thereof. In some embodiments, the Tn7-type transposase complex comprises at least one or at least two, three polypeptides comprising a sequence having at least 80% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111, or a variant thereof.In some embodiments, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222, or a variant thereof. In some embodiments, the engineered guide polynucleotide comprises a sequence having at least 80% sequence identity to any one of the non-degenerate nucleotides of SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, or 123-140, or a variant thereof. In some embodiments, the left-side recombinase sequence comprises a sequence having at least 80% identity to SEQ ID NOs: 9, 11, 36-38, 76, and 78, or a variant thereof. In some embodiments, the right-side recombinase sequence comprises a sequence having at least 80% identity to SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93, or a variant thereof. In some embodiments, the class 2 V-type Cas effector and the Tn7-type transposase complex are encoded by a polynucleotide sequence comprising less than about 10 kilobases.

[0005] In some aspects, the disclosure provides a method for translocating a cargo nucleotide sequence into a target nucleic acid site comprising a target nucleotide sequence, comprising expressing in a cell or introducing into a cell a system of any of the aspects or embodiments described herein.

[0006] In some aspects, the disclosure provides a method for transposing a cargo nucleotide sequence into a target nucleic acid site, comprising contacting a first double-stranded nucleic acid comprising the cargo nucleotide sequence with a Cas effector complex comprising a class 2 V-type Cas effector and an engineered guide polynucleotide configured to hybridize to the target nucleic acid site, a Tn7-type transposase complex configured to bind to the Cas effector complex, the Tn7-type transposase complex comprising a TnsB subunit, and a second double-stranded nucleic acid comprising the target nucleic acid site. In some embodiments, the cargo nucleotide sequence is flanked by a left transposase recognition sequence and a right transposase recognition sequence. In some embodiments, the system further comprises a PAM sequence compatible with the nuclease flanking the target nucleic acid site. In some embodiments, the PAM sequence is located 3' of the target nucleic acid site. In some embodiments, the engineered guide polynucleotide is configured to bind to the class 2 V-type Cas effector. In some embodiments, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least 80% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220, or a variant thereof. In some embodiments, the TnsB subunit comprises a polypeptide comprising a sequence having at least 80% identity to SEQ ID NOs: 2, 13, 17, and 65, or a variant thereof. In some embodiments, the Tn7-type transposase complex comprises at least one or at least two polypeptides comprising a sequence having at least 80% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111, or a variant thereof. In some embodiments, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222, or a variant thereof.In some embodiments, the left recombinase sequence comprises a sequence having at least 80% identity to SEQ ID NOs: 9, 11, 36-38, 76, and 78, or a variant thereof. In some embodiments, the right recombinase sequence comprises a sequence having at least 80% identity to SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93, or a variant thereof. In some embodiments, the class 2 V-type Cas effector and the Tn7-type transposase complex are encoded by a polynucleotide sequence comprising less than about 10 kilobases.

[0007] In some aspects, the disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site, the system comprising: a first double-stranded nucleic acid comprising a cargo nucleotide sequence configured to interact with a Tn7-type transposase complex; a Cas effector complex comprising a class 2 V-type Cas effector and an engineered guide polynucleotide configured to hybridize to the target nucleic acid site; and a Tn7-type transposase complex configured to bind to the Cas effector complex, the Tn7-type transposase complex comprising TnsB, TnsC, and TniQ components. and a Tn7-type transposase complex, wherein (a) the class 2 V-type Cas effector comprises a polypeptide having a sequence with at least 80% sequence identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220, or a variant thereof, or b) the Tn7-type transposase complex comprises a TnsB, TnsC, or TniQ component having a sequence with at least 80% sequence identity to any one of SEQ ID NOs: 2-4, 13-15, 17-19, 65-67, or 109-111, or a variant thereof. In some embodiments, the transposase complex is non-covalently linked to the Cas effector complex. In some embodiments, the transposase complex is covalently linked to the Cas effector complex. In some embodiments, the transposase complex is fused to the Cas effector complex in a single polypeptide. In some embodiments, the class 2 V-type Cas effector comprises a polypeptide having a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220, or a variant thereof. In some embodiments, the Tn7-type transposase complex comprises a TnsB, TnsC, or TniQ component having a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 2-4, 13-15, 17-19, 65-67, or 109-111, or a variant thereof.In some embodiments, the class 2 V-type Cas effector is a Cas12k effector. In some embodiments, the cargo nucleotide sequence is flanked by a left transposase recognition sequence and a right transposase recognition sequence. In some embodiments, the system further comprises a second double-stranded nucleic acid comprising the target nucleic acid site. In some embodiments, the system further comprises a PAM sequence compatible with the nuclease flanking the target nucleic acid site. In some embodiments, the PAM sequence is located 5' or 3' of the target nucleic acid site. In some embodiments, the PAM sequence comprises SEQ ID NO: 31. In some embodiments, the engineered guide polynucleotide is configured to bind to the class 2 V-type Cas effector. In some embodiments, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222, or a variant thereof. In some embodiments, the engineered guide polynucleotide comprises a sequence having at least 80% sequence identity to a non-degenerate nucleotide of any one of SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, or 123-140, or a variant thereof. In some embodiments, the left-side recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78, or a variant thereof. In some embodiments, the right-side recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some embodiments, the class 2 V-type Cas effector and the Tn7-type transposase complex are encoded by a polynucleotide sequence comprising less than about 10 kilobases.In some embodiments, (a) the class 2 V-type Cas effector comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1, 81, 82, 83, or 85, or a variant thereof; (b) the left recombinase sequence comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 9, 11, 36, 37, or 38, or a variant thereof; or (c) the right recombinase sequence comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 8, 39, 40, 41, 42, 43, 44, or 93, or a variant thereof. (d) the engineered guide polynucleotide comprises (i) a sequence having at least 80% sequence identity to at least 46-80 nucleotides of SEQ ID NO:6, or (ii) a sequence having at least 80% identity to the non-degenerate nucleotides of any one of SEQ ID NOs:5, 45-63, 68-75, 96-103, or 123-140, or a variant thereof; (e) the TnsB, TnsC, and TniQ components comprise a polypeptide having a sequence having at least 80% identity to SEQ ID NOs:2-4, or a variant thereof; or (f) the PAM sequence comprises SEQ ID NO:31.In some embodiments, (a) the class 2 V-type Cas effector comprises a sequence having at least 80% sequence identity to SEQ ID NO: 12, or a variant thereof; (b) the left recombinase sequence comprises a sequence having at least 80% sequence identity to SEQ ID NO: 76, or a variant thereof; (c) the right recombinase sequence comprises a sequence having at least 80% sequence identity to SEQ ID NO: 77, or a variant thereof; (d) the engineered guide polynucleotide comprises (i) a sequence having at least 80% sequence identity to at least 46-80 nucleotides of SEQ ID NO: 32 or 104, or (ii) a sequence having at least 80% identity to any one of the non-degenerate nucleotides of SEQ ID NO: 107 or 102, or a variant thereof; or (e) the TnsB, TnsC, and TniQ components comprise a polypeptide having a sequence having at least 80% identity to SEQ ID NO: 13-15, or a variant thereof. In some embodiments, (a) the class 2 V-type Cas effector comprises a sequence having at least 80% sequence identity to SEQ ID NO: 16, or a variant thereof; (b) the left recombinase sequence comprises a sequence having at least 80% sequence identity to SEQ ID NO: 78, or a variant thereof; (c) the right recombinase sequence comprises a sequence having at least 80% sequence identity to SEQ ID NO: 79, or a variant thereof; (d) the engineered guide polynucleotide comprises (i) a sequence having at least 80% sequence identity to at least 46-80 nucleotides of SEQ ID NO: 33 or 105, or (ii) a sequence having at least 80% identity to any one of the non-degenerate nucleotides of SEQ ID NO: 108 or 103, or a variant thereof; or (e) the TnsB, TnsC, and TniQ components comprise a polypeptide having a sequence having at least 80% identity to SEQ ID NO: 17-19, or a variant thereof.

[0008] In some aspects, the disclosure provides an engineered nuclease system comprising: an endonuclease comprising a RuvC domain, wherein the endonuclease is derived from an uncultured microorganism, and wherein the endonuclease is a class 2 VK-type Cas effector having at least 80% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220, or a variant thereof; and an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, and wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a target nucleic acid sequence. In some embodiments, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222, or a variant thereof. In some embodiments, the engineered guide polynucleotide comprises a sequence having at least 80% identity to a non-degenerate nucleotide of any one of SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, or 123-140, or a variant thereof. In some embodiments, the system further comprises a PAM sequence compatible with the nuclease adjacent to the target nucleic acid site. In some embodiments, the PAM sequence is located 5' of the target nucleic acid site. In some embodiments, the PAM sequence comprises SEQ ID NO: 31.In some embodiments, (a) the class 2 VK-type Cas effector comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1, 81, 82, 83, or 85, or a variant thereof; (b) the left recombinase sequence comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 9, 11, 36, 37, or 38, or a variant thereof; or (c) the right recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 8, 39, 40, 41, 42, 43, 44, or 93, or a variant thereof. (d) the engineered guide polynucleotide comprises (i) a sequence having at least 80% sequence identity to at least 46-80 nucleotides of SEQ ID NO:6, or a variant thereof, or (ii) a sequence having at least 80% identity to the non-degenerate nucleotides of any one of SEQ ID NOs:5, 45-63, 68-75, 96-103, or 123-140, or a variant thereof; (e) the TnsB, TnsC, and TniQ components comprise polypeptides having at least 80% identity to SEQ ID NOs:2-4, or a variant thereof; or (f) the PAM sequence comprises SEQ ID NO:31.

[0009] In some aspects, the disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, the system comprising: a Cas effector complex comprising a class 2 V-type Cas effector, a small prokaryotic ribosomal protein subunit S15, and an engineered guide polynucleotide configured to hybridize to the target nucleic acid site; a Tn7-type transposase complex configured to bind to the Cas effector complex and comprising TnsB, TnsC, and TniQ components; and a double-stranded nucleic acid configured to interact with the Tn7-type transposase complex and comprising a cargo nucleotide sequence. In some embodiments, the Cas effector complex is non-covalently bound to the Tn7-type transposase complex. In some embodiments, the Cas effector complex is covalently bound to the Tn7-type transposase complex. In some embodiments, the Cas effector complex is fused to the Tn7-type transposase complex.

[0010] In some embodiments, the cargo nucleotide sequence is flanked by left and right transposase recognition sequences recognized by a Tn7-type transposase complex. In some embodiments, the left recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some embodiments, the right recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93.

[0011] In some embodiments, the target nucleic acid comprises a PAM sequence that is compatible with a Cas effector complex. In some embodiments, the PAM sequence comprises SEQ ID NO: 31. In some embodiments, the PAM sequence is located about 50 to about 70 base pairs from the target nucleic acid site. In some embodiments, the PAM sequence is located 3' to the target nucleic acid site. In some embodiments, the PAM sequence is located 5' to the target nucleic acid site.

[0012] In some embodiments, the class 2 V-type Cas effector is a Cas12k effector. In some embodiments, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least 80% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least 90% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least 80% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the TnsB component comprises a polypeptide comprising a sequence having at least 80% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide having a sequence having at least 90% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide having a sequence having at least 90% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least 80% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least 90% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising any one of the sequences of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.

[0013] In some embodiments, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the engineered guide polynucleotide comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, and 123-140.

[0014] In some embodiments, the small prokaryotic ribosomal protein subunit S15 comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 187-189. In some embodiments, the small prokaryotic ribosomal protein subunit S15 is encoded by a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 181-183.

[0015] In some embodiments, the class 2 V-type Cas effector and the Tn7-type transposase complex are encoded by a polynucleotide sequence comprising less than about 10 kilobases.

[0016] In some aspects, the disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, the system comprising a Cas effector complex comprising a class 2 V-type Cas effector and an engineered guide polynucleotide configured to hybridize to the target nucleic acid site, the Cas effector complex comprising a polypeptide comprising a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. The present invention provides a system comprising: a Tn7-type transposase complex configured to bind to a Cas effector complex and comprising TnsB, TnsC, and TniQ components, wherein the TnsB, TnsC, and TniQ components comprise a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 2-4, 13-15, 17-19, 65-67, and 109-111; and a double-stranded nucleic acid configured to interact with the Tn7-type transposase complex and comprising a cargo nucleotide sequence.

[0017] In some embodiments, the Cas effector complex is non-covalently linked to the Tn7-type transposase complex. In some embodiments, the Cas effector complex is covalently linked to the Tn7-type transposase complex. In some embodiments, the Cas effector complex is fused to the Tn7-type transposase complex.

[0018] In some embodiments, the cargo nucleotide sequence is flanked by left and right transposase recognition sequences recognized by a Tn7-type transposase complex. In some embodiments, the left recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some embodiments, the right recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93.

[0019] In some embodiments, the target nucleic acid comprises a PAM sequence that is compatible with a Cas effector complex. In some embodiments, the PAM sequence comprises SEQ ID NO: 31. In some embodiments, the PAM sequence is located about 50 to about 70 base pairs from the target nucleic acid site. In some embodiments, the PAM sequence is located 3' to the target nucleic acid site. In some embodiments, the PAM sequence is located 5' to the target nucleic acid site.

[0020] In some embodiments, the class 2 V-type Cas effector is a Cas12k effector. In some embodiments, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least 90% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence of any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220.

[0021] In some embodiments, the TnsB, TnsC, or TniQ component comprises a sequence having at least 90% sequence identity to any one of SEQ ID NOs: 2-4, 13-15, 17-19, 65-67, and 109-111. In some embodiments, the TnsB, TnsC, or TniQ component comprises a sequence of any one of SEQ ID NOs: 2-4, 13-15, 17-19, 65-67, and 109-111.

[0022] In some embodiments, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the engineered guide polynucleotide comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, and 123-140.

[0023] In some embodiments, the engineered nuclease comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 187-189. In some embodiments, the small prokaryotic ribosomal protein subunit S15 is encoded by a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 181-183.

[0024] In some embodiments, the class 2 V-type Cas effector and the Tn7-type transposase complex are encoded by a polynucleotide sequence comprising less than about 10 kilobases.

[0025] In some aspects, the disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, the system comprising a Cas effector complex comprising a class 2 V-type Cas effector comprising a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1, 81, 82, 83, and 85, and an engineered guide polynucleotide comprising at least 80% identity to any one of SEQ ID NOs: 5, 6, 45-63, 68-75, 96-103, and 123-140, and a Tn7-type transposer configured to bind to the Cas effector complex and comprising TnsB, TnsC, and TniQ components. and a double-stranded nucleic acid configured to interact with the Tn7-type transposase complex, the double-stranded nucleic acid comprising, in 5' to 3' order, a left recombinase sequence that comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 9, 11, 36, 37, and 38, a cargo nucleotide sequence, and a right recombinase sequence that comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 8, 39-44, and 93.

[0026] In some aspects, the disclosure provides a system for translocating a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, the system comprising a Cas effector complex configured to hybridize to the target nucleic acid site, the Cas effector complex comprising a class 2 V-type Cas effector comprising a sequence having at least 80% sequence identity to SEQ ID NO: 12, and an engineered guide polynucleotide comprising at least 80% identity to any one of SEQ ID NOs: 32, 102, 104, and 107, and a Cas effector complex configured to bind to the Cas effector complex, the Cas effector complex comprising TnsB, TnsC, and TniQ components. The present invention provides a system comprising: a Tn7-type transposase complex, wherein the TnsB, TnsC, or TniQ component comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 13-15; and a double-stranded nucleic acid configured to interact with the Tn7-type transposase complex, the double-stranded nucleic acid comprising, in 5' to 3' order, a left recombinase sequence comprising a sequence having at least 80% sequence identity to SEQ ID NO: 76, a cargo nucleotide sequence, and a right recombinase sequence comprising a sequence having at least 80% identity to SEQ ID NO: 77.

[0027] In some aspects, the disclosure provides a system for translocating a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, the system comprising a Cas effector complex configured to hybridize to the target nucleic acid site, the Cas effector complex comprising a class 2 V-type Cas effector comprising a sequence having at least 80% sequence identity to SEQ ID NO: 16, and an engineered guide polynucleotide comprising at least 80% identity to any one of SEQ ID NOs: 33, 103, 105, and 108, and an engineered guide polynucleotide configured to bind to the Cas effector complex, the engineered guide polynucleotide comprising TnsB, TnsC, and TniQ components. The present invention provides a system comprising: a Tn7-type transposase complex, wherein a TnsB, TnsC, or TniQ component comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 17-19; and a double-stranded nucleic acid configured to interact with the Tn7-type transposase complex, the double-stranded nucleic acid comprising, in 5' to 3' order, a left recombinase sequence comprising a sequence having at least 80% sequence identity to SEQ ID NO: 78, a cargo nucleotide sequence, and a right recombinase sequence comprising a sequence having at least 80% identity to SEQ ID NO: 79.

[0028] In some aspects, the disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, the system comprising a Cas effector complex comprising a class 2 V-type Cas effector comprising a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1, 81, 82, 83, and 85, and an engineered guide polynucleotide comprising at least 80% identity to any one of SEQ ID NOs: 5, 6, 45-63, 68-75, 96-103, and 123-140, and a TnSEQ ID NO: 1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90 The system includes a Tn7-type transposase complex, the TnsB, TnsC, and TniQ components comprising a sequence having at least 80% identity to any one of SEQ ID NOs:2-4, and a double-stranded nucleic acid configured to interact with the Tn7-type transposase complex, the double-stranded nucleic acid comprising, in 5' to 3' order, a left recombinase sequence comprising a sequence having at least 80% sequence identity to SEQ ID NOs:9, 11, 36, 37, and 38, a cargo nucleotide sequence, and a right recombinase sequence comprising a sequence having at least 80% identity to SEQ ID NOs:8, 39-44, and 93.

[0029] In some embodiments, the system further comprises a PAM sequence compatible with a Cas effector complex. In some embodiments, the PAM sequence comprises SEQ ID NO:31.

[0030] In some embodiments, the PAM sequence is located about 50 to about 70 base pairs from the target nucleic acid site. In some embodiments, the PAM sequence is located 3' to the target nucleic acid site. In some embodiments, the PAM sequence is located 5' to the target nucleic acid site.

[0031] In some embodiments, the Cas effector complex further comprises a small prokaryotic ribosomal protein subunit S15. In some embodiments, the small prokaryotic ribosomal protein subunit S15 comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 187-189. In some embodiments, the small prokaryotic ribosomal protein subunit S15 is encoded by a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 181-183.

[0032] In some aspects, the disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, the system comprising: a Cas effector complex comprising a class 2 V-type Cas effector, a small prokaryotic ribosomal protein subunit S15, and an engineered guide polynucleotide, the Cas effector complex comprising the engineered guide polynucleotide configured to hybridize to the target nucleic acid site; a Tn7-type transposase complex operably linked to the Cas effector complex, the Tn7-type transposase complex comprising TnsB, TnsC, and TniQ components; and a double-stranded nucleic acid comprising, in 5' to 3' order, a left recombinase recognition sequence, a cargo nucleotide sequence, and a right recombinase recognition sequence, wherein the left recombinase recognition sequence and the right recombinase recognition sequence are recognized by the Tn7-type transposase complex.

[0033] In some embodiments, the left recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some embodiments, the right recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93.

[0034] In some embodiments, the target nucleic acid comprises a PAM sequence that is compatible with a Cas effector complex. In some embodiments, the PAM sequence comprises SEQ ID NO: 31. In some embodiments, the PAM sequence is located about 50 to about 70 base pairs from the target nucleic acid site. In some embodiments, the PAM sequence is located 3' to the target nucleic acid site. In some embodiments, the PAM sequence is located 5' to the target nucleic acid site.

[0035] In some embodiments, the class 2 V-type Cas effector is a Cas12k effector. In some embodiments, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least 80% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least 90% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least 90% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220.

[0036] In some embodiments, the TnsB subunit comprises a polypeptide having a sequence having at least 80% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB subunit comprises a polypeptide having a sequence having at least 90% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB subunit comprises a polypeptide having a sequence having any one of SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least 80% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least 90% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least 90% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.

[0037] In some embodiments, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the engineered guide polynucleotide comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, and 123-140.

[0038] In some embodiments, the engineered nuclease comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 187-189. In some embodiments, the small prokaryotic ribosomal protein subunit S15 is encoded by a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 181-183.

[0039] In some embodiments, the class 2 V-type Cas effector and the Tn7-type transposase complex are encoded by a polynucleotide sequence comprising less than about 10 kilobases.

[0040] In some aspects, the disclosure provides an engineered nuclease system comprising an endonuclease comprising a RuvC domain, the endonuclease being derived from an uncultured microorganism and being a class 2 VK-type Cas effector having at least 80% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220; and an engineered guide RNA configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to a target nucleic acid sequence.

[0041] In some embodiments, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the engineered guide polynucleotide comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, and 123-140.

[0042] In some aspects, the present disclosure provides a method for translocating a cargo nucleotide sequence into a target nucleic acid site comprising introducing a system of the present disclosure into a cell.

[0043] In some aspects, the disclosure provides a cell comprising the system of the disclosure. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is an immortalized cell. In some embodiments, the cell is an insect cell. In some embodiments, the cell is a yeast cell. In some embodiments, the cell is a plant cell. In some embodiments, the cell is a fungal cell. In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is an A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cells, HT1080, HepG2, Huh7, K562, primary cells, or derivatives thereof. In some embodiments, the cell is an engineered cell. In some embodiments, the cells are stable cells.

[0044] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, in which only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modification in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive. [Brief description of the drawings]

[0045] The novel features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings (also referred to herein as "Figure" and "FIG.").

[0046] [Figure 1] 1 depicts exemplary organizations of different classes and types of CRISPR / Cas loci. [Diagram 2]1 depicts the structure of the natural class 2 type II crRNA / tracrRNA pair compared to the hybrid sgRNA in which the crRNA and tracrRNA are joined. [Diagram 3] Two pathways found in Tn7 and Tn7-like elements are depicted. [Figure 4A] The genomic context of the MG64 family's type V Tn7 CAST is depicted. Figure 4A depicts that the MG64-1 CAST system contains a CRISPR array (CRISPR repeats), a type V nuclease, and three predicted transposase protein sequences. A tracrRNA was predicted within the intergenic region between the CAST effector array and the CRISPR array. Bottom: Multiple sequence alignment of the catalytic domain of transposase TnsB. Catalytic residues are indicated by boxes. Figure 4B depicts that two transposon ends were predicted for the MG64-1 CAST system. [Figure 4B] The genomic context of the MG64 family's type V Tn7 CAST is depicted. Figure 4A depicts that the MG64-1 CAST system contains a CRISPR array (CRISPR repeats), a type V nuclease, and three predicted transposase protein sequences. A tracrRNA was predicted within the intergenic region between the CAST effector array and the CRISPR array. Bottom: Multiple sequence alignment of the catalytic domain of transposase TnsB. Catalytic residues are indicated by boxes. Figure 4B depicts that two transposon ends were predicted for the MG64-1 CAST system. [Diagram 5] The predicted structure of the corresponding sgRNA of the CAST system described herein is depicted. Panel A of Figure 5 (left) shows the predicted MG64-1 tracrRNA and crRNA double-stranded complex in the repeat-antirepeat stem. The loop was cleaved and a GAAA tetraloop was added to the stem-loop structure to produce the designed sgRNA shown in Panel B of Figure 5 (right). [Figure 6]Depicts the results of transposition reactions targeted with a plasmid library consisting of NNNNNNNN at the 5' of the target spacer sequence. Reaction #1 shows the presence of the targeted library, #2 shows the presence of the donor fragment in both transposition reactions, and #3-5 show the sg-specific PCR bands corresponding to the correct transposition reactions. [Figure 7A] 7A depicts the results of Sanger sequencing. FIG. 7A shows Sanger sequencing of the donor-target junction on the left end (LE) of the transposon in a LE transposition reaction closer to the PAM. The expected sequence is at the top of the panel, with the predicted transposition event 61 bp away from the PAM. The top chromatogram is the sequencing result starting within the donor fragment. A clear signal is seen on the far right up to the donor / target junction (dotted line). This shows a mixture of transposition products. The bottom chromatogram of the panel is sequencing from the target to the donor / target junction. The signal from the left is a clear signal up to the junction. [Figure 7B] 7A-7C depict the results of Sanger sequencing. FIG. 7B shows Sanger sequencing of the donor-target junction on the right end (RE) of the transposon in the product of the LE closer to the PAM. The expected sequence is at the top of the panel, with the predicted transposition event 61 bp away from the PAM. The top chromatogram is the sequencing result starting within the donor fragment. A clear signal is seen to the left all the way to the donor / target junction (dotted line). [Figure 7C] Figure 7C depicts the results of Sanger sequencing. Figure 7C is a contiguity map of the PAM library. [Figure 7D] 7A-7D are SeqLogo analysis of NGS of LE events closer to the PAM, showing a very strong preference for NGTN at the PAM motif. [Figure 8]Figure 1 depicts a phylogenetic gene tree of Cas12k effector sequences, inferred from a multiple sequence alignment of 64 Cas12k sequences recovered here (orange and black branches) and 229 reference Cas12k sequences from public databases (grey branch). The orange branch indicates Cas12k effectors confirmed to be associated with CAST transposon components. [Figure 9] MG64 family CRISPR repeat alignments are shown. The Cas12k CAST CRISPR repeat contains the conserved motif 5'-GNNGGNNTGAAAG-3'. In MG64-1, a short repeat-anti-repeat (RAR) within the CRISPR repeat motif aligns with the tracrRNA. The MG64 RAR motif appears to define the start and end of the tracrRNA (5' end: RAR1 (TTTC), 3' end: RAR2 (CCNNC)). [Figure 10A] Depicts the secondary structure predicted from folding of CRISPR repeats+tracrRNA for the MG64 system. [Figure 10B] Depicts the secondary structure predicted from folding of CRISPR repeats+tracrRNA for the MG64 system. [Figure 11A] Depicts the MG64-3 CRISPR locus. The tracrRNA is encoded upstream from the CRISPR array, while the transposon ends are encoded downstream (inner black box). Sequences corresponding to partial 3' CRISPR repeats and partial spacers are encoded within the transposon (outer box). Self-matching spacers are encoded outside the transposon ends. [Figure 11B]Depicting tracrRNA sequence alignments for various CASTs provided herein. Alignment of tracrRNA sequences shows regions of conservation. In particular, the sequence "TGCTTTC" (upper box) at sequence positions 92-98 may be important for sgRNA tertiary structure and for non-contiguous repeat-anti-repeat pairing with the crRNA. The hairpin "CYCC(n6)GGRG" (lower box) at positions 265-278 may be important for function, such as by positioning downstream sequences for crRNA pairing. [Figure 12A] 1 depicts the predicted structure of MG64-1 sgRNA. [Figure 12B] 1 depicts the predicted structure of MG64-3 sgRNA. [Figure 12C] 1 depicts the predicted structure of the MG64-5 sgRNA. [Figure 13A] Depicts PCR data demonstrating that MG64-1 is active with sgRNA v2-1. Using the protocol described for in vitro target integrase activity, the effector protein and its TnsB, TnsC, and TniQ proteins were expressed in an in vitro transcription / translation system. After translation, target DNA, cargo DNA, and sgRNA were added in the reaction buffer. Integration was assayed by PCR across the target / donor junction. Depicts a diagram illustrating the potential orientations of integrated donor DNA. PCR reactions 3, 4, 5, and 6 represent each integrated ligation product depending on the orientation in which the donor was integrated at the target site. [Figure 13B]Depicts PCR data demonstrating that MG64-1 is active with sgRNA v2-1. Using the protocol described for in vitro target integrase activity, the effector protein and its TnsB, TnsC, and TniQ proteins were expressed in an in vitro transcription / translation system. After translation, target DNA, cargo DNA, and sgRNA were added in the reaction buffer. Integration was assayed by PCR across the target / donor junction. Depicts gel images of PCR4 of transposition (detecting RE junction to donor) showing lane 1) apo (no sgRNA), lane 2) with sgRNA 1, and lane 3) with sgRNA v2-1. [Figure 13C] Depicts PCR data demonstrating that MG64-1 is active with sgRNA v2-1. Using the protocol described for in vitro target integrase activity, the effector protein and its TnsB, TnsC, and TniQ proteins were expressed in an in vitro transcription / translation system. After translation, target DNA, cargo DNA, and sgRNA were added in reaction buffer. Integration was assayed by PCR across the target / donor junction. Depicts gel images of PCR5 of transposition (detecting LE junction to donor) showing lane 1) apo (no sgRNA), lane 2) with sgRNA 1, and lane 3) with sgRNA v2-1. [Figure 14] Depicts PCR reaction 5 (LE proximal to the PAM, top half of the plot) and PCR reaction 4 (RE distal to the PAM, bottom half of the plot) plotted on sequence and distance from the PAM of MG64-1. Analysis of integration windows indicates that 95% of integrations occurring at the spacer PAM site fall within a 10 bp window 58-68 nucleotides away from the PAM. The difference in integration distance between distal and proximal frequencies reflects integration site overlap, i.e., a 3-5 base pair overlap as a result of the staggered nuclease activity of the transposase during integration. [Figure 15]Depicts the results of colony PCR screening of transposition efficiency. After incubation, 18 colony forming units (CFU) were visible on the plate, 8 on plate A (no IPTG, lane labeled as A) and 10 on plate B (100 μM IPTG during recovery, lane labeled as B). All 18 were analyzed by colony PCR, which gave rise to product bands (arrows) indicating a successful transposition reaction. [Figure 16] Sequencing results of selected colony PCR products are depicted, confirming that they represent transposition events since they span the junction between the LE and PAM at the engineered target site within the lacZ gene. The minimal LE sequence is shown in blue at the top of the screen (Min LE), while the target and PAM are shown in grey. Some sequence variation is observed in the PCR products, but this variation is expected considering that insertions can occur at variable distances upstream of the PAM. [Figure 17]FIG. 17 depicts the results of testing engineered single guides for 64-1 transposition activity. Black boxes are lanes not relevant to this experiment. Panel A of FIG. 17 depicts a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = sgRNA v1-1, lane 4 = sgRNA v1-2, lane 5 = sgRNA v1-3. Panel B of FIG. 17 depicts a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = sgRNA v1-1, lane 4 = sgRNA v1-2, lane 5 = sgRNA v1-3. Panel C of Figure 17 depicts a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = sgRNA v1-4, lane 4 = sgRNA v1-6, lane 5 = sgRNA v1-7, lane 6 = sgRNA v1-8, lane 7 = sgRNA v1-9. Panel D of Figure 17 depicts a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = sgRNA v1-4, lane 4 = sgRNA v1-6, lane 5 = sgRNA v1-7, lane 6 = sgRNA v1-8, lane 7 = sgRNA v1-9. Panel E of Figure 17 depicts a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = sgRNA v1-5, lane 4 = skip, lane 5 = sgRNA v1-10. Panel F of Figure 17 depicts a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = sgRNA v1-5, lane 4 = skip, lane 5 = sgRNA v1-10.Panel G of Figure 17 depicts a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = sgRNAv1-17, lane 4 = sgRNA v1-18, lane 5 = skip, lane 6 = sgRNA v1-19, lane 7 = skip, lane 8 = sgRNA v1-20. Panel H of Figure 17 depicts a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = sgRNAv1-17, lane 4 = sgRNA v1-18, lane 5 = skip, lane 6 = sgRNA v1-19, lane 7 = skip, lane 8 = sgRNA v1-20. [Figure 18]Figure 18 shows the results of testing engineered LE and RE on 64-1 transposition activity. Black boxes are lanes not relevant to this experiment. Panel A of Figure 18 depicts a gel image of PCR4 of transposition (detecting RE junction to donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = LE 86bp, lane 4 = LE 105bp, lane 5 = RE 196bp, lane 6 = RE 242bp, lane 7 = RE internal deletion 50, lane 8 = RE internal deletion 81. Panel B of Figure 18 depicts a gel image of PCR 5 of the transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = LE 86 bp, lane 4 = LE 105 bp, lane 5 = RE 196 bp, lane 6 = RE 242 bp, lane 7 = RE internal deletion 50, lane 8 = RE internal deletion 81. Panel C of Figure 18 depicts a gel image of PCR 4 of the transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = RE internal deletions 81 and 178 bp, lane 4 = skip, lane 5 = RE internal deletions 81 and 196 bp, lane 6 = skip, lane 7 = RE internal deletions 81 and 212 bp, lane 8 = skip. Panel D of Figure 18 shows a gel image of PCR 5 of the transposition (detecting the LE junction to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = RE internal deletion 81 and 178bp, lane 4 = skip, lane 5 = RE internal deletion 81 and 196bp, lane 6 = skip, lane 7 = RE internal deletion 81 and 212bp, lane 8 = skip. Panel E of Figure 18 depicts a gel image of PCR 4 of the transposition (detecting the RE junction to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = RE internal deletion 81 and 178bp + LE 68bp, lane 4 = RE internal deletion 81 and 178bp + LE 86bp, lane 5 = skip, lane 6 = RE internal deletion 81 and 178bp + LE 105bp, lane 7 = skip.Panel F of Figure 18 depicts a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = RE internal deletions 81 and 178bp + LE 68bp, lane 4 = RE internal deletions 81 and 178bp + LE 86bp, lane 5 = skip, lane 6 = RE internal deletions 81 and 178bp + LE 105bp, lane 7 = skip. Panel G of Figure 18 depicts a gel image of PCR 6 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = 0bp overhang, lane 4 = 1bp overhang, lane 5 = 2bp overhang, lane 6 = 3bp overhang, lane 7 = 5bp overhang, lane 8 = 10bp overhang. [Figure 19]FIG. 19 depicts the results of testing engineered CAST components containing NLS for transposition activity. Black boxes are lanes not relevant to this experiment. Panel A of FIG. 19 depicts a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1=apo (no sgRNA), lane 2=holo (+sgRNA), lane 3=skip, lane 4=skip, lane 5=skip, lane 6=NLS-TnsB, lane 7=skip, lane 8=TnsB-NLS. Panel B of FIG. 19 depicts a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1=apo (no sgRNA), lane 2=holo (+sgRNA), lane 3=skip, lane 4=skip, lane 5=skip, lane 6=NLS-TnsB, lane 7=skip, lane 8=TnsB-NLS. Panel C of Figure 19 depicts a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = skip, lane 4 = skip, lane 5 = skip, lane 6 = NLS-TniQ, lane 7 = skip, lane 8 = TniQ-NLS. Panel D of Figure 19 depicts a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = skip, lane 4 = skip, lane 5 = skip, lane 6 = NLS-TniQ, lane 7 = skip, lane 8 = TniQ-NLS. Panel E of Figure 19 depicts a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = skip, lane 4 = skip, lane 5 = NLS-Cas12k, lane 6 = Cas12k-NLS, lane 7 = NLS-TnsC, lane 8 = TnsC-NLS. Panel F of Figure 19 depicts a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = skip, lane 4 = skip, lane 5 = NLS-Cas12k, lane 6 = Cas12k-NLS, lane 7 = NLS-TnsC, lane 8 = TnsC-NLS.Panel G of Figure 19 depicts a gel image of PCR 4 of transposition (detecting RE junction to donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = NLS-HA-TnsC, lane 4 = NLS-TnsC-FLAG, lane 5 = NLS-TnsC-HA, lane 6 = NLS-TnsC-Myc, lane 7 = NLS-FLAG-TnsC, lane 8 = NLS-Myc-TnsC. Panel H of Figure 19 depicts a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = NLS-HA-TnsC, lane 4 = NLS-TnsC-FLAG, lane 5 = NLS-TnsC-HA, lane 6 = NLS-TnsC-Myc, lane 7 = NLS-FLAG-TnsC, lane 8 = NLS-Myc-TnsC. Panel I of Figure 19 depicts a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = Cas 2x NLS apo (no sgRNA), lane 4 = Cas 2x NLS holo (+sgRNA). Panel J of Figure 19 depicts a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = Cas 2x NLS apo (no sgRNA), lane 4 = Cas 2x NLS holo (+sgRNA). [Figure 20]The engineered CAST-NLS acting as a single suite is depicted. All lanes have Cas12k-NLS and NLS-TniQ, TnsB, TnsC, and sgRNA unless otherwise noted. Panel A of Figure 20 depicts a gel image of PCR4 of transposition (detecting RE junction to donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = NLS-TnsB, lane 4 = TnsB-NLS, lane 5 = NLS-TnsB and NLS-TnsC, lane 6 = TnsB-NLS and NLS-TnsC. Panel B of Figure 20 shows a gel image of PCR 5 of the transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = NLS-TnsB, lane 4 = TnsB-NLS, lane 5 = NLS-TnsB and NLS-TnsC, lane 6 = TnsB-NLS and NLS-TnsC. [Figure 21]The results of testing Cas effector and TniQ protein fusions on transposition activity are depicted. Panel A of FIG. 21 depicts a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo with Cas-TniQ fusion (no sgRNA), lane 2 = holo with Cas-TniQ fusion (+sgRNA), lane 3 = apo with TniQ-Cas fusion (no sgRNA), lane 4 = holo with TniQ-Cas fusion (+sgRNA). Panel B of FIG. 21 depicts a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo with Cas-TniQ fusion (no sgRNA), lane 2 = holo with Cas-TniQ fusion (+sgRNA), lane 3 = apo with TniQ-Cas fusion (no sgRNA), lane 4 = holo with TniQ-Cas fusion (+sgRNA). Panel C of Figure 21 depicts a gel image of PCR4 of transposition (detecting RE junction to donor): lane 1 = apo with TniQ-Cas fusion (no sgRNA), lane 2 = holo with TniQ-Cas fusion (+sgRNA), lane 3 = holo-Cas alone, lane 4 = apo with TniQ-48 linker-Cas fusion (no sgRNA), lane 5 = holo with TniQ-48 linker-Cas fusion (+sgRNA), lane 6 = apo with TniQ-68 linker-Cas fusion (no sgRNA), lane 7 = holo with TniQ-68 linker-Cas fusion (+sgRNA), lane 8 = holo with TniQ-72 linker-Cas fusion (+sgRNA). Panel D of Figure 21 depicts a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo with TniQ-Cas fusion (no sgRNA), lane 2 = holo with TniQ-Cas fusion (+sgRNA), lane 3 = holo-Cas alone, lane 4 = apo with TniQ-48 linker-Cas fusion (no sgRNA), lane 5 = holo with TniQ-48 linker-Cas fusion (+sgRNA), lane 6 = apo with TniQ-68 linker-Cas fusion (no sgRNA), lane 7 = holo with TniQ-68 linker-Cas fusion (+sgRNA), lane 8 = holo with TniQ-72 linker-Cas fusion (+sgRNA).Panel E of Figure 21 depicts a gel image of PCR4 of transposition (detecting RE junction to donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = apo (no sgRNA) with NLS-TniQ-Cas-NLS fusion, lane 4 = holo (+sgRNA) with NLS-TniQ-Cas-NLS fusion, lane 5 = apo (no sgRNA) with NLS-TniQ-77 linker-Cas-NLS fusion, lane 6 = holo (+sgRNA) with NLS-TniQ-77 linker-Cas-NLS fusion. Panel F of Figure 21 depicts a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = apo (no sgRNA) with NLS-TniQ-Cas-NLS fusion, lane 4 = holo (+sgRNA) with NLS-TniQ-Cas-NLS fusion, lane 5 = apo (no sgRNA) with NLS-TniQ-77 linker-Cas-NLS fusion, lane 6 = holo (+sgRNA) with NLS-TniQ-77 linker-Cas-NLS fusion. Panel G of Figure 21 depicts a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = NLS-TniQ-Cas-NLS apo (no sgRNA), lane 4 = NLS-TniQ-Cas-NLS holo (+sgRNA), lane 5 = Cas-NLS-P2A-NLS-TniQ apo (no sgRNA), lane 6 = Cas-NLS-P2A-NLS-TniQ holo (+sgRNA). Panel H of Figure 21 depicts a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = NLS-TniQ-Cas-NLS apo (no sgRNA), lane 4 = NLS-TniQ-Cas-NLS holo (+sgRNA), lane 5 = Cas-NLS-P2A-NLS-TniQ apo (no sgRNA), lane 6 = Cas-NLS-P2A-NLS-TniQ holo (+sgRNA). [Figure 22]Expression of TnsB and TnsC in human cells, followed by cell fractionation and the results of an in vitro transposition reaction are depicted. Panel A of Figure 22 depicts a gel image of PCR4 (detecting RE junctions to the donor) of transposition: lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = holo (+sgRNA) with untreated (no TnsB) cytoplasm, lane 4 = holo (+sgRNA) with untreated nucleoplasm, lane 5 = holo (+sgRNA) with NLS-TnsB cytoplasm, lane 6 = holo (+sgRNA) with NLS-TnsB nucleoplasm, lane 7 = holo (+sgRNA) with TnsB-NLS cytoplasm, lane 8 = holo (+sgRNA) with TnsB-NLS nucleoplasm, lane 9 = holo (+sgRNA) with NLS-TniQ cytoplasm, lane 10 = holo (+sgRNA) with NLS-TniQ nucleoplasm. Panel B of Figure 22 depicts a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = holo (+sgRNA) with untreated (no TnsB) cytoplasm, lane 4 = holo (+sgRNA) with untreated nucleoplasm, lane 5 = holo (+sgRNA) with NLS-TnsB cytoplasm, lane 6 = holo (+sgRNA) with NLS-TnsB nucleoplasm, lane 7 = holo (+sgRNA) with TnsB-NLS cytoplasm, lane 8 = holo (+sgRNA) with TnsB-NLS nucleoplasm, lane 9 = holo (+sgRNA) with NLS-TniQ cytoplasm, lane 10 = holo (+sgRNA) with NLS-TniQ nucleoplasm. Panel C of Figure 22 depicts a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = holo (+sgRNA) without TnsC, lane 4 = holo (+sgRNA) with untreated (no TnsC) cytoplasm, lane 5 = holo (+sgRNA) with untreated nucleoplasm, lane 6 = holo (+sgRNA) with NLS-HA-TnsC cytoplasm, lane 7 = holo (+sgRNA) with NLS-HA-TnsC nucleoplasm, lane 8 = holo (+sgRNA) with TnsC-NLS cytoplasm, lane 9 = holo (+sgRNA) with TnsC-NLS nucleoplasm.Panel D of Figure 22 depicts a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = holo (+sgRNA) without TnsC, lane 4 = holo (+sgRNA) with untreated (no TnsC) cytoplasm, lane 5 = holo (+sgRNA) with untreated nucleoplasm, lane 6 = holo (+sgRNA) with NLS-HA-TnsC cytoplasm, lane 7 = holo (+sgRNA) with NLS-HA-TnsC nucleoplasm, lane 8 = holo (+sgRNA) with TnsC-NLS cytoplasm, lane 9 = holo (+sgRNA) with TnsC-NLS nucleoplasm. Panel E of Figure 22 depicts a gel image of PCR4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = apo (no sgRNA) NLS-TnsB-IRES-NLS-TnsC cytoplasm, lane 4 = holo (+sgRNA) NLS-TnsB-IRES-NLS-TnsC cytoplasm, lane 5 = apo (no sgRNA) NLS-TnsB-IRES-NLS-TnsC nucleoplasm, lane Lane 6 = holo (+sgRNA) NLS-TnsB-IRES-NLS-TnsC nucleoplasm, lane 7 = apo (no sgRNA) TnsB-NLS-IRES-NLS-TnsC cytoplasm, lane 8 = holo (+sgRNA) TnsB-NLS-IRES-NLS-TnsC cytoplasm, lane 9 = apo (no sgRNA) TnsB-NLS-IRES-NLS-TnsC nucleoplasm, lane 10 = holo (+sgRNA) TnsB-NLS-IRES-NLS-TnsC nucleoplasm.Panel F of Figure 22 depicts a gel image of PCR 5 of transposition (detecting LE junctions to donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = apo (no sgRNA) NLS-TnsB-IRES-NLS-TnsC cytoplasmic, lane 4 = holo (+sgRNA) NLS-TnsB-IRES-NLS-TnsC cytoplasmic, lane 5 = apo (no sgRNA) NLS-TnsB-IRES-NLS-TnsC lane 5 = holo (+sgRNA) TnsB-NLS-IRES-NLS-TnsC nucleoplasm, lane 6 = holo (+sgRNA) TnsB-NLS-IRES-NLS-TnsC nucleoplasm, lane 7 = apo (no sgRNA) TnsB-NLS-IRES-NLS-TnsC cytoplasm, lane 8 = holo (+sgRNA) TnsB-NLS-IRES-NLS-TnsC cytoplasm, lane 9 = apo (no sgRNA) TnsB-NLS-IRES-NLS-TnsC nucleoplasm, lane 10 = holo (+sgRNA) TnsB-NLS-IRES-NLS-TnsC nucleoplasm. [Figure 23]Figure 23 depicts the expression of Cas12k and TniQ combined constructs in human cells, followed by the results of an in vitro transposition assay. Panel A of Figure 23 depicts a gel image of PCR5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = Cas-NLS holo (+sgRNA) cytoplasm, lane 4 = Cas-NLS holo (+sgRNA) nucleoplasm, lane 5 = Cas-NLS holo (+sgRNA) nucleoplasm + additional sgRNA, lane 6 = Cas-NLS-P2A-NLS-TniQ holo (+sgRNA) cytoplasm, lane 7 = Cas-NLS-P2A-NLS-TniQ holo (+sgRNA) cytoplasm, lane 8 = Cas-NLS-P2A-NLS-TniQ holo (+sgRNA) cytoplasm + additional sgRNA. Panel B of Figure 23 depicts a gel image of PCR4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = apo (no sgRNA) Cas-NLS-P2A-NLS-TniQ cytoplasm, lane 4 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ cytoplasm, lane 5 = apo (no sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm, lane 6 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm, lane 7 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm + additional holo Cas-NLS, lane 8 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm + NLS-TniQ.Panel C of Figure 23 depicts a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = apo (no sgRNA) Cas-NLS-P2A-NLS-TniQ cytoplasm, lane 4 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ cytoplasm, lane 5 = apo (no sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm, lane 6 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm, lane 7 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm + additional holo Cas-NLS, lane 8 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm + NLS-TniQ. Panel D of Figure 23 depicts a gel image of PCR4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = apo (no sgRNA) NLS-TniQ-Cas-NLS cytoplasm, lane 4 = holo (+sgRNA) NLS-TniQ-Cas-NLS cytoplasm, lane 5 = apo (no sgRNA) NLS-TniQ-Cas-NLS nucleoplasm, lane 6 = holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm, lane 7 = holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm + additional holo-Cas-NLS, lane 8 = holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm + NLS-TniQ. Panel E of Figure 23 depicts a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = apo (no sgRNA) NLS-TniQ-Cas-NLS cytoplasm, lane 4 = holo (+sgRNA) NLS-TniQ-Cas-NLS cytoplasm, lane 5 = apo (no sgRNA) NLS-TniQ-Cas-NLS nucleoplasm, lane 6 = holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm, lane 7 = holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm + additional holo-Cas-NLS, lane 8 = holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm + NLS-TniQ.Panel F of Figure 23 shows a gel image of PCR4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ cytoplasm, lane 4 = holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ cytoplasm, lane 5 = apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm, lane 6 = apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional PURExpress, lane 7 = apo (no sgRNA) Cas-NLS Lane 5 = apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional Cas-NLS, lane 6 = apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + NLS-TniQ, lane 9 = holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm, lane 10 = holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional PURExpress, lane 11 = holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional Cas-NLS, lane 12 = holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + NLS-TniQ.Panel G of Figure 23 depicts a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ cytoplasmic, lane 4 = holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ cytoplasmic, lane 5 = apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasmic, lane 6 = apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasmic + additional PURExpress, lane 7 = apo (no sgRNA) Cas- Lane 8 = apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + NLS-TniQ, lane 9 = holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm, lane 10 = holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional PURExpress, lane 11 = holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional Cas-NLS, lane 12 = holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + NLS-TniQ. [Figure 24] 1 depicts the results of an electrophoretic mobility shift assay (EMSA) of 64-1 TnsB and its LE DNA sequence. The EMSA results confirm binding and TnsB recognition. TnsB protein was expressed in an in vitro transcription / translation system, incubated with FAM-labeled DNA containing the LE sequence, and then separated on a native 5% TBE gel. Binding is observed as an upward shift in the labeled band. Multiple TnsB binding sites lead to multiple shifts in the EMSA. Lane 1: FAM-labeled DNA only. Lane 2: FAM DNA + in vitro transcription / translation system (no TnsB protein). Lane 3: FAM DNA + TnsB. [Figure 25A]Depicting Cas12k effector diversity. Figure 25A depicts the Cas12k CAST genomic context. The transposon is characterized by terminal inverted repeats (TIRs, light orange bars), Tn7-like transposon genes (colored arrows), the death effector Cas12k (orange arrow), tracrRNA (pink half arrow), and the CRISPR array. A "TAAA" target site duplication (TDS) was observed adjacent to the TIR. Center panel: Inset of MG64-1 non-coding region showing tracrRNA, pseudo-repeats and self-targeting spacers, CRISPR array, and transposon left end TIR. Bottom panel: Multiple alignment of pseudo-repeats and self-targeting spacers in the group of CAST homologs. Figure 25B depicts an unrooted phylogenetic tree of Cas12k effectors. The Cas12k effectors recovered in this study are shown as orange (confirmed transposons in the genome) and black branches, while the reference Cas12k sequences are shown in grey. The reference sequences ShCas12k and AcCas12k are indicated by red arrows. [Figure 25B]Depicting Cas12k effector diversity. Figure 25A depicts the Cas12k CAST genomic context. The transposon is characterized by terminal inverted repeats (TIRs, light orange bars), Tn7-like transposon genes (colored arrows), the death effector Cas12k (orange arrow), tracrRNA (pink half arrow), and the CRISPR array. A "TAAA" target site duplication (TDS) was observed adjacent to the TIR. Center panel: Inset of MG64-1 non-coding region showing tracrRNA, pseudo-repeats and self-targeting spacers, CRISPR array, and transposon left end TIR. Bottom panel: Multiple alignment of pseudo-repeats and self-targeting spacers in the group of CAST homologs. Figure 25B depicts an unrooted phylogenetic tree of Cas12k effectors. The Cas12k effectors recovered in this study are shown as orange (confirmed transposons in the genome) and black branches, while the reference Cas12k sequences are shown in grey. The reference sequences ShCas12k and AcCas12k are indicated by red arrows. [Figure 26A] The multiple sequence alignments of the CAST right (Figure 26A) and left (Figure 26B) ends are depicted. The transposon terminal inversion motif "TGTNNA" is highlighted in a box. [Figure 26B] The multiple sequence alignments of the CAST right (Figure 26A) and left (Figure 26B) ends are depicted. The transposon terminal inversion motif "TGTNNA" is highlighted in a box. [Figure 27] Depicts an alignment of Cas12k CAST tracrRNA sequences showing regions of sequence and structural conservation. In particular, the sequence "TGCTTTC" at sequence positions 88-92 may be important for sgRNA tertiary structure and for non-contiguous repeat-antirepeat pairing with the crRNA. The hairpin "CYCC(n6)GGRG" at positions 279-294 may be important for function, potentially positioning downstream sequences for crRNA pairing. [Figure 28]1 depicts single guide RNA folding of active MG64-1, MG64-2, and MG64-6 CAST systems. The active engineered sgRNA of MG64-1 is also shown. [Figure 29A] 1 depicts in vitro screening of CAST rearrangements using a PAM library. 2 depicts the screening setup for in vitro PAM determination. [Figure 29B] 1 depicts in vitro screening of CAST transposition using a PAM library. 2 depicts a schematic diagram of junction PCR for detection of transposition products. [Figure 30A] Depicts the translocation junctions of MG64-1 CAST (left lane) and MG64-6 CAST (right lane) amplified by PCR. [Figure 30B] Depicts the SeqLogo representation of the detected PAMs of MG64-1 (top). [Figure 30C] FIG. 1 depicts integration frequency plotted by distance over proximal and distal distances of MG64-1. [Diagram 31] Depicting single-guide RNA engineering of 64-1. Deletion of an approximately 130 bp to 190 bp region (green and cyan sections of the structure) generated an sgRNA that induced a strong transposition reaction (green bar on heatmap). [Diagram 32] Depicts MG64-2 sgRNA cross-reactivity with PAMs for MG64-1 and MG64-2 sgRNA + MG64-1 effector combinations. [Diagram 33] Depicts single guide RNA cleavage in the coding DNA of MG64-2 sgRNA in sequence diagram and secondary structure prediction model. The deleted regions and cleavage in the sequence diagram are shown as bars (del1, del2, del3, del4, del5, and del6). The deleted regions in ellipses in the secondary structure prediction model indicate the tested cleavage (deletion 1, 2, 3, 4 across pseudoknot, 5, and 6). [Diagram 34]29 depicts data demonstrating that engineered MG64-2 sgRNAs are active in the MG64-1 CAST system. PCR reactions represent each possible integration junction or negative control (FIG. 29B). Products of successful integration are highlighted by arrows. Boxed lanes are not relevant to this experiment. [Diagram 35] Depicting the MG64-2 sgRNA split-guide design, the sgRNA fragments were synthesized separately and then reannealed before being tested in transposition experiments. [Diagram 36] 29 depicts data demonstrating that split MG64-2 sgRNAs are active in the MG64-1 CAST system. PCR reactions represent each possible integration junction or negative control (FIG. 29B). Products of successful integration are highlighted by arrows. Boxed lanes are not relevant to this experiment. [Figure 37] 1 depicts data demonstrating that LE and RE minimization maintained transposition activity of the system. [Figure 38A] Figure 38A depicts the results of E. coli integration with MG64-1. Figure 38B depicts a schematic of the introduction of the CAST system into E. coli. Figure 38B depicts NGS data showing an editing efficiency of over 80%. Figure 38C depicts off-target analysis showing that off-target integration was not detected in more than 1% of all total transposition events. [Figure 38B] Figure 38A depicts the results of E. coli integration with MG64-1. Figure 38B depicts a schematic of the introduction of the CAST system into E. coli. Figure 38B depicts NGS data showing an editing efficiency of over 80%. Figure 38C depicts off-target analysis showing that off-target integration was not detected in more than 1% of all total transposition events. [Figure 38C] Figure 38A depicts the results of E. coli integration with MG64-1. Figure 38B depicts a schematic of the introduction of the CAST system into E. coli. Figure 38B depicts NGS data showing an editing efficiency of over 80%. Figure 38C depicts off-target analysis showing that off-target integration was not detected in more than 1% of all total transposition events. [Figure 39] Depict the local insertion rates of various endogenous loci in the E. coli genome. [Figure 40A] Figure 40 depicts the results of multigene coordinate targeting. Figure 40A depicts the local insertion frequency at endogenous locus and engineered locus, respectively. Figure 40B depicts the relative insertion frequency of on-target insertion at endogenous locus, on-target insertion at engineered locus, and off-target insertion. The integration at both loci combined accounts for more than 95% of all integrations that occurred on the genome. [Figure 40B] Figure 40 depicts the results of multigene coordinate targeting. Figure 40A depicts the local insertion frequency at endogenous locus and engineered locus, respectively. Figure 40B depicts the relative insertion frequency of on-target insertion at endogenous locus, on-target insertion at engineered locus, and off-target insertion. The integration at both loci combined accounts for more than 95% of all integrations that occurred on the genome. [Diagram 41] 1 depicts Sanger sequencing data of integrated PCR products demonstrating that MG64-1 is active in vitro. The reaction is of the RE donor target product, and the point where sequencing stops matching the donor DNA is when the junction occurs (darker bar below the sequencing peak). [Figure 42A] Figure 1 shows a schematic of serial dilution of target DNA for in vitro transposition experiments. CAST components are expressed in PureExpress and added to reactions with in vitro transcribed sgRNA and donor plasmid. Target plasmid DNA is added at decreasing concentrations and tested for transposition experiments. When the minimum amount of target DNA is determined, transposition reactions are assayed by adding increasing amounts of human genomic DNA. [Figure 42B] A diagram of PCR amplification of transposition reactions is shown. An 8N PAM plasmid library (8N-target, Rxn#1) is targeted with the CAST system to integrate donor DNA (Rxn#2). Upon successful integration, junction PCR reactions are performed with primers to amplify four putative integration reactions based on the orientation of cargo integration (Rxn#3, #4, #5, and #6). [Figure 42C]Illustrates PCR reaction products from an in vitro transposition assay using serial dilutions of target plasmid DNA. Target, donor, and reactions #3, #4, #5, and #6 correspond to the PCR integration products shown in Figure 42B. [Fig.42D] PCR reaction products from an in vitro transposition assay using a fixed amount of target plasmid DNA (0.5 ng) while increasing amounts of human genomic DNA were added to increase the search space. Target, donor, and reactions #3, #4, #5, and #6 correspond to the PCR integration products shown in Figure 42B. [Figure 43A] A schematic of a transposition reaction across a high copy element is shown. The target PCR product spans the wild type target element when assayed with CAST protein and an sgRNA targeting one of multiple aligned targets. Integration can occur in either the forward orientation, reverse orientation, or both. The forward transposition product is assayed by junction PCR, which amplifies a region encompassing the LE of the donor DNA to the 5' end of the target site (Fwd PCR). The reverse junction reaction assays a region encompassing the LE of the donor DNA to the 3' end of the target element (Rev PCR). [Figure 43B] 43A shows PCR reaction products from an in vitro transposition assay at 15 target sites (guides) in the line 1 3' element in human genomic DNA. The targets and Fwd and Rev PCR of the reactions correspond to the PCR integration products shown in FIG. 43A. [Figure 43C] 43A-43C show PCR reaction products from an in vitro transposition assay at 15 target sites (guides) in the SVA element in human genomic DNA. The targets and Fwd and Rev PCR of the reactions correspond to the PCR integration products shown in FIG. 43A. Bands highlighted with arrows indicate successful targeted integration. [Fig. 43D] Figure 43 shows PCR reaction products from an in vitro transposition assay at 15 target sites (guides) in HERV elements in human genomic DNA. The targets and Fwd and Rev PCR of the reactions correspond to the PCR integration products shown in Figure 43A. Bands highlighted with arrows indicate successful targeted integration. [Figure 43E] Line 1 shows Sanger sequencing of Fwd PCR integration products at multiple target sites in the 3' element. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. [Figure 43F] Line 1 shows Sanger sequencing of Rev PCR integration products at multiple target sites in the 3' element. The point where the sequencing trace stops matching the target DNA (gray vertical bar) is the point where integration occurs. [Figure 43G] Sanger sequencing of the Fwd PCR integration product at SVA target site 3 is shown. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. [Fig. 43H] Sanger sequencing of the Fwd PCR product at HERV target site 5 is shown. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. [Diagram 44] Shown are PCR reaction products from an in vitro transposition assay at line 1 target sites 12 and 15 in human genomic DNA with functional domains. The targets and Fwd and Rev PCR of the reactions correspond to the PCR integration products shown in Figure 42A. The bands highlighted with arrows indicate successful targeted integration. [Figure 45A]Illustrated are in vitro transposition experiments with CAST, S15, NLS-S15, and S15-NLS expressed from eukaryotic transcription / translation reactions. Figure 45A shows an in vitro transposition reaction with MG64-1 CAST and S15. Wheat germ extract-expressed CAST components promote transposition without the addition of S15, albeit at a slower rate (faint band highlighted with an arrow). Addition of PURExpress reagent (spent PUREx) increases transposition efficiency as indicated by the intensity of the band in Rxn#5 (PURExpress reagent contains S15). Independent addition of S15 and S15-NLS translated from wheat germ extract reactions increases transposition efficiency by MG64-1 in vitro compared to other conditions tested (intense band highlighted with an arrow). Figure 45B shows an in vitro reaction of transposition using the NLS-S15 construct. Addition of PURExpress reagent increases in vitro transposition (lane 3) compared to CAST components only conditions (lane 2). The NLS-S15 construct did not improve transposition (lanes 4-5). The boxed Rxn#5 represents the expected band if transposition activity is detected. [Figure 45B]Illustrated are in vitro transposition experiments with CAST, S15, NLS-S15, and S15-NLS expressed from eukaryotic transcription / translation reactions. Figure 45A shows an in vitro transposition reaction with MG64-1 CAST and S15. Wheat germ extract-expressed CAST components promote transposition without the addition of S15, albeit at a slower rate (faint band highlighted with an arrow). Addition of PURExpress reagent (spent PUREx) increases transposition efficiency as indicated by the intensity of the band in Rxn#5 (PURExpress reagent contains S15). Independent addition of S15 and S15-NLS translated from wheat germ extract reactions increases transposition efficiency by MG64-1 in vitro compared to other conditions tested (intense band highlighted with an arrow). Figure 45B shows an in vitro reaction of transposition using the NLS-S15 construct. Addition of PURExpress reagent increases in vitro transposition (lane 3) compared to CAST components only conditions (lane 2). The NLS-S15 construct did not improve transposition (lanes 4-5). The boxed Rxn#5 represents the expected band if transposition activity is detected. [Figure 46A]Schematic diagram of fusion plasmids for cell transposition. Figure 46A: Two targeting complex plasmids and one donor plasmid are assembled for high copy element line 1, targets 8, 12, and 15, and SVA target 3. Figure 46B shows cell transposition into high copy elements using H1core-TniQ or HMGN1-TniQ in line 1 targets 8, 12, 15, and SVA target 3. Arrows indicate amplified transposition junction reactions in either forward (Fwd PCR) or reverse (Rev PCR) orientation of transposition. Mock controls represent reactions without targeting or donor plasmid. Figure 46C shows Sanger sequencing of PCR integration product Fwd PCR at line 1 3' target site 8. Integration was mediated by MG64-1 with NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46D shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 8. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46E shows Sanger sequencing of the PCR integration product Rev PCR at line 1 3' target site 12. Integration was mediated by MG64-1 with an NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46F shows Sanger sequencing of the PCR integration product Rev PCR at line 1 3' target site 12. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46G shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 15. Integration was mediated by MG64-1 carrying the NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs.Figure 46H shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 15. Integration was mediated by MG64-1 carrying the NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. [Figure 46B]Schematic diagram of fusion plasmids for cell transposition. Figure 46A: Two targeting complex plasmids and one donor plasmid are assembled for high copy element line 1, targets 8, 12, and 15, and SVA target 3. Figure 46B shows cell transposition into high copy elements using H1core-TniQ or HMGN1-TniQ in line 1 targets 8, 12, 15, and SVA target 3. Arrows indicate amplified transposition junction reactions in either forward (Fwd PCR) or reverse (Rev PCR) orientation of transposition. Mock controls represent reactions without targeting or donor plasmid. Figure 46C shows Sanger sequencing of PCR integration product Fwd PCR at line 1 3' target site 8. Integration was mediated by MG64-1 with NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46D shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 8. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46E shows Sanger sequencing of the PCR integration product Rev PCR at line 1 3' target site 12. Integration was mediated by MG64-1 with an NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46F shows Sanger sequencing of the PCR integration product Rev PCR at line 1 3' target site 12. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46G shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 15. Integration was mediated by MG64-1 carrying the NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs.Figure 46H shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 15. Integration was mediated by MG64-1 carrying the NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. [Figure 46C]Schematic diagram of fusion plasmids for cell transposition. Figure 46A: Two targeting complex plasmids and one donor plasmid are assembled for high copy element line 1, targets 8, 12, and 15, and SVA target 3. Figure 46B shows cell transposition into high copy elements using H1core-TniQ or HMGN1-TniQ in line 1 targets 8, 12, 15, and SVA target 3. Arrows indicate amplified transposition junction reactions in either forward (Fwd PCR) or reverse (Rev PCR) orientation of transposition. Mock controls represent reactions without targeting or donor plasmid. Figure 46C shows Sanger sequencing of PCR integration product Fwd PCR at line 1 3' target site 8. Integration was mediated by MG64-1 with NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46D shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 8. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46E shows Sanger sequencing of the PCR integration product Rev PCR at line 1 3' target site 12. Integration was mediated by MG64-1 with an NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46F shows Sanger sequencing of the PCR integration product Rev PCR at line 1 3' target site 12. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46G shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 15. Integration was mediated by MG64-1 carrying the NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs.Figure 46H shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 15. Integration was mediated by MG64-1 carrying the NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. [Figure 46D]Schematic diagram of fusion plasmids for cell transposition. Figure 46A: Two targeting complex plasmids and one donor plasmid are assembled for high copy element line 1, targets 8, 12, and 15, and SVA target 3. Figure 46B shows cell transposition into high copy elements using H1core-TniQ or HMGN1-TniQ in line 1 targets 8, 12, 15, and SVA target 3. Arrows indicate amplified transposition junction reactions in either forward (Fwd PCR) or reverse (Rev PCR) orientation of transposition. Mock controls represent reactions without targeting or donor plasmid. Figure 46C shows Sanger sequencing of PCR integration product Fwd PCR at line 1 3' target site 8. Integration was mediated by MG64-1 with NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46D shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 8. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46E shows Sanger sequencing of the PCR integration product Rev PCR at line 1 3' target site 12. Integration was mediated by MG64-1 with an NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46F shows Sanger sequencing of the PCR integration product Rev PCR at line 1 3' target site 12. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46G shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 15. Integration was mediated by MG64-1 carrying the NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs.Figure 46H shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 15. Integration was mediated by MG64-1 carrying the NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. [Figure 46E]Schematic diagram of fusion plasmids for cell transposition. Figure 46A: Two targeting complex plasmids and one donor plasmid are assembled for high copy element line 1, targets 8, 12, and 15, and SVA target 3. Figure 46B shows cell transposition into high copy elements using H1core-TniQ or HMGN1-TniQ in line 1 targets 8, 12, 15, and SVA target 3. Arrows indicate amplified transposition junction reactions in either forward (Fwd PCR) or reverse (Rev PCR) orientation of transposition. Mock controls represent reactions without targeting or donor plasmid. Figure 46C shows Sanger sequencing of PCR integration product Fwd PCR at line 1 3' target site 8. Integration was mediated by MG64-1 with NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46D shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 8. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46E shows Sanger sequencing of the PCR integration product Rev PCR at line 1 3' target site 12. Integration was mediated by MG64-1 with an NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46F shows Sanger sequencing of the PCR integration product Rev PCR at line 1 3' target site 12. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46G shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 15. Integration was mediated by MG64-1 carrying the NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs.Figure 46H shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 15. Integration was mediated by MG64-1 carrying the NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. [Figure 46F]Schematic diagram of fusion plasmids for cell transposition. Figure 46A: Two targeting complex plasmids and one donor plasmid are assembled for high copy element line 1, targets 8, 12, and 15, and SVA target 3. Figure 46B shows cell transposition into high copy elements using H1core-TniQ or HMGN1-TniQ in line 1 targets 8, 12, 15, and SVA target 3. Arrows indicate amplified transposition junction reactions in either forward (Fwd PCR) or reverse (Rev PCR) orientation of transposition. Mock controls represent reactions without targeting or donor plasmid. Figure 46C shows Sanger sequencing of PCR integration product Fwd PCR at line 1 3' target site 8. Integration was mediated by MG64-1 with NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46D shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 8. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46E shows Sanger sequencing of the PCR integration product Rev PCR at line 1 3' target site 12. Integration was mediated by MG64-1 with an NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46F shows Sanger sequencing of the PCR integration product Rev PCR at line 1 3' target site 12. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46G shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 15. Integration was mediated by MG64-1 carrying the NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs.Figure 46H shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 15. Integration was mediated by MG64-1 carrying the NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. [Figure 46G]Schematic diagram of fusion plasmids for cell transposition. Figure 46A: Two targeting complex plasmids and one donor plasmid are assembled for high copy element line 1, targets 8, 12, and 15, and SVA target 3. Figure 46B shows cell transposition into high copy elements using H1core-TniQ or HMGN1-TniQ in line 1 targets 8, 12, 15, and SVA target 3. Arrows indicate amplified transposition junction reactions in either forward (Fwd PCR) or reverse (Rev PCR) orientation of transposition. Mock controls represent reactions without targeting or donor plasmid. Figure 46C shows Sanger sequencing of PCR integration product Fwd PCR at line 1 3' target site 8. Integration was mediated by MG64-1 with NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46D shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 8. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46E shows Sanger sequencing of the PCR integration product Rev PCR at line 1 3' target site 12. Integration was mediated by MG64-1 with an NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46F shows Sanger sequencing of the PCR integration product Rev PCR at line 1 3' target site 12. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46G shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 15. Integration was mediated by MG64-1 carrying the NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs.Figure 46H shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 15. Integration was mediated by MG64-1 carrying the NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. [Fig. 46H]Schematic diagram of fusion plasmids for cell transposition. Figure 46A: Two targeting complex plasmids and one donor plasmid are assembled for high copy element line 1, targets 8, 12, and 15, and SVA target 3. Figure 46B shows cell transposition into high copy elements using H1core-TniQ or HMGN1-TniQ in line 1 targets 8, 12, 15, and SVA target 3. Arrows indicate amplified transposition junction reactions in either forward (Fwd PCR) or reverse (Rev PCR) orientation of transposition. Mock controls represent reactions without targeting or donor plasmid. Figure 46C shows Sanger sequencing of PCR integration product Fwd PCR at line 1 3' target site 8. Integration was mediated by MG64-1 with NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46D shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 8. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46E shows Sanger sequencing of the PCR integration product Rev PCR at line 1 3' target site 12. Integration was mediated by MG64-1 with an NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46F shows Sanger sequencing of the PCR integration product Rev PCR at line 1 3' target site 12. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. Figure 46G shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 15. Integration was mediated by MG64-1 carrying the NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs.Figure 46H shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 15. Integration was mediated by MG64-1 carrying the NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. [Figure 47A] Immunofluorescence staining for localization of Cas12k CAST components in human cells. Figure 47A: Top row: detection of TnsB localization, middle row: detection of Cas12k localization, bottom row: detection of TnsC localization. Images showed that MG64-1 Cas12k and TnsB are localized in the nucleus of mammalian cells, while TnsC is localized in the cytoplasm. Cas12k CAST protein was tagged with HA tag. Anti-HA antibody was used for protein detection. DAPI was used to stain DNA (nucleus). Figure 47B: Top and bottom rows: detection of TniQ localization. Images showed that MG64-1 TniQ is localized in the nucleus of mammalian cells. CAST protein was tagged with HA tag. Anti-HA antibody was used for protein detection. DAPI was used to stain DNA (nucleus). Figure 47C: All rows: detection of TnsC colocalization with TniQ. The images show that while some TnsC may remain in the cytoplasm, here it co-localizes with TniQ in the nucleus. CAST protein was tagged with an HA tag. Anti-HA antibody was used for protein detection. DAPI was used to stain DNA (nucleus). Figure 47D: Both rows: Cas12k, TnsB, TnsC, and TniQ co-delivered to HEK293T cells localize in the nucleus. CAST protein was tagged with an HA tag. Anti-HA antibody was used for protein detection. DAPI was used to stain DNA (nucleus). [Figure 47B]Immunofluorescence staining for localization of Cas12k CAST components in human cells. Figure 47A: Top row: detection of TnsB localization, middle row: detection of Cas12k localization, bottom row: detection of TnsC localization. Images showed that MG64-1 Cas12k and TnsB are localized in the nucleus of mammalian cells, while TnsC is localized in the cytoplasm. Cas12k CAST protein was tagged with HA tag. Anti-HA antibody was used for protein detection. DAPI was used to stain DNA (nucleus). Figure 47B: Top and bottom rows: detection of TniQ localization. Images showed that MG64-1 TniQ is localized in the nucleus of mammalian cells. CAST protein was tagged with HA tag. Anti-HA antibody was used for protein detection. DAPI was used to stain DNA (nucleus). Figure 47C: All rows: detection of TnsC colocalization with TniQ. The images show that while some TnsC may remain in the cytoplasm, here it co-localizes with TniQ in the nucleus. CAST protein was tagged with an HA tag. Anti-HA antibody was used for protein detection. DAPI was used to stain DNA (nucleus). Figure 47D: Both rows: Cas12k, TnsB, TnsC, and TniQ co-delivered to HEK293T cells localize in the nucleus. CAST protein was tagged with an HA tag. Anti-HA antibody was used for protein detection. DAPI was used to stain DNA (nucleus). [Figure 47C]Immunofluorescence staining for localization of Cas12k CAST components in human cells. Figure 47A: Top row: detection of TnsB localization, middle row: detection of Cas12k localization, bottom row: detection of TnsC localization. Images showed that MG64-1 Cas12k and TnsB are localized in the nucleus of mammalian cells, while TnsC is localized in the cytoplasm. Cas12k CAST protein was tagged with HA tag. Anti-HA antibody was used for protein detection. DAPI was used to stain DNA (nucleus). Figure 47B: Top and bottom rows: detection of TniQ localization. Images showed that MG64-1 TniQ is localized in the nucleus of mammalian cells. CAST protein was tagged with HA tag. Anti-HA antibody was used for protein detection. DAPI was used to stain DNA (nucleus). Figure 47C: All rows: detection of TnsC colocalization with TniQ. The images show that while some TnsC may remain in the cytoplasm, here it co-localizes with TniQ in the nucleus. CAST protein was tagged with an HA tag. Anti-HA antibody was used for protein detection. DAPI was used to stain DNA (nucleus). Figure 47D: Both rows: Cas12k, TnsB, TnsC, and TniQ co-delivered to HEK293T cells localize in the nucleus. CAST protein was tagged with an HA tag. Anti-HA antibody was used for protein detection. DAPI was used to stain DNA (nucleus). [Figure 47D]Immunofluorescence staining for localization of Cas12k CAST components in human cells. Figure 47A: Top row: detection of TnsB localization, middle row: detection of Cas12k localization, bottom row: detection of TnsC localization. Images showed that MG64-1 Cas12k and TnsB are localized in the nucleus of mammalian cells, while TnsC is localized in the cytoplasm. Cas12k CAST protein was tagged with HA tag. Anti-HA antibody was used for protein detection. DAPI was used to stain DNA (nucleus). Figure 47B: Top and bottom rows: detection of TniQ localization. Images showed that MG64-1 TniQ is localized in the nucleus of mammalian cells. CAST protein was tagged with HA tag. Anti-HA antibody was used for protein detection. DAPI was used to stain DNA (nucleus). Figure 47C: All rows: detection of TnsC colocalization with TniQ. The images show that while some TnsC may remain in the cytoplasm, here it co-localizes with TniQ in the nucleus. CAST protein was tagged with an HA tag. Anti-HA antibody was used for protein detection. DAPI was used to stain DNA (nucleus). Figure 47D: Both rows: Cas12k, TnsB, TnsC, and TniQ co-delivered to HEK293T cells localize in the nucleus. CAST protein was tagged with an HA tag. Anti-HA antibody was used for protein detection. DAPI was used to stain DNA (nucleus). [Figure 48A] Depicts in vitro screening of MG64-1 Cas12k CAST transposition. Figure 48A: Schematic of constructs used for MG64-1 holocomplex purification. Figure 48B: Schematic of junction PCR for detection of transposition products. Target substrate with 5'PAM followed by protospacer (target, Rxn#1) is targeted with the CAST system to integrate cargo DNA (Rxn#2). Upon successful integration, junction PCR reactions are performed with primers to amplify the four putative integration reactions based on the orientation of cargo integration. [Figure 48B]Depicts in vitro screening of MG64-1 Cas12k CAST transposition. Figure 48A: Schematic of constructs used for MG64-1 holocomplex purification. Figure 48B: Schematic of junction PCR for detection of transposition products. Target substrate with 5'PAM followed by protospacer (target, Rxn#1) is targeted with the CAST system to integrate cargo DNA (Rxn#2). Upon successful integration, junction PCR reactions are performed with primers to amplify the four putative integration reactions based on the orientation of cargo integration. [Figure 48C] Depicting MG64-1 protein purification. Figure 48C: Fractions collected during 2 L-scale purification of MG64-1 holocomplex run on an unstained denaturing PAGE gel. Figure 48D: Chromatogram of size exclusion chromatography (SEC) performed on MG64-1 holocomplex. The peak centered at 29.3 mL (peak 1) was used for in vitro activity assays. [Figure 48D] Depicting MG64-1 protein purification. Figure 48C: Fractions collected during 2 L-scale purification of MG64-1 holocomplex run on an unstained denaturing PAGE gel. Figure 48D: Chromatogram of size exclusion chromatography (SEC) performed on MG64-1 holocomplex. The peak centered at 29.3 mL (peak 1) was used for in vitro activity assays. [Figure 49A]Figure 49 depicts in vitro transposition using peak 1 recovered holo complex supplemented with TnT expression components. Lane L) ladder, lane 1) TnT expression CAST components apo condition (-sgRNA), lane 2) TnT expression CAST components holo condition (+sgRNA), lane 3) purified peak 1 complemented with TnT CAST components without additional supplementation of Cas12k (-TnT Cas12k), lane 4) purified peak 1 complemented with TnT CAST components without additional supplementation of TnsC (-TnT TnsC), lane 5) peak 1 complemented with TnT CAST components without additional supplementation of TniQ (-TnT TniQ), lane 6) peak 1 complemented with TnT CAST components without additional supplementation of S15 (-TnT S15). Figure 49B depicts Sanger sequencing of lane 3, lane 4, lane 5, and lane 6 from both pDonor and target orientations of the amplified LE-PAM target-donor junctions. The vertical line defines the predicted translocation junction for MG64-1 in the reference sequence. Degradation of the signal from either direction results in multiple signals reflected in the PCR amplification. [Figure 49B] Figure 49 depicts in vitro transposition using peak 1 recovered holo complex supplemented with TnT expression components. Lane L) ladder, lane 1) TnT expression CAST components apo condition (-sgRNA), lane 2) TnT expression CAST components holo condition (+sgRNA), lane 3) purified peak 1 complemented with TnT CAST components without additional supplementation of Cas12k (-TnT Cas12k), lane 4) purified peak 1 complemented with TnT CAST components without additional supplementation of TnsC (-TnT TnsC), lane 5) peak 1 complemented with TnT CAST components without additional supplementation of TniQ (-TnT TniQ), lane 6) peak 1 complemented with TnT CAST components without additional supplementation of S15 (-TnT S15). Figure 49B depicts Sanger sequencing of lane 3, lane 4, lane 5, and lane 6 from both pDonor and target orientations of the amplified LE-PAM target-donor junctions. The vertical line defines the predicted translocation junction for MG64-1 in the reference sequence. Degradation of the signal from either direction results in multiple signals reflected in the PCR amplification. [Figure 50]Depicting the identification of ribosomal protein S15 homologues in cyanobacterial genomic fragments. Candidate sequences from the same sample in which MG64-1 was recovered are highlighted with dark circles. The reference S15 from E. coli is indicated with an arrow.

[0047] Brief Description of the Sequence Listing The Sequence Listing submitted herewith provides exemplary polynucleotide and polypeptide sequences for use in the methods, compositions, and systems according to the present disclosure. Below are exemplary descriptions of the sequences therein.

[0048] MG64 SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220 show the full-length peptide sequences of the MG64 Cas effector.

[0049] SEQ ID NOs: 2-4, 13-15, 17-19, 65-67, and 109-111 show peptide sequences of MG64 translocation proteins that may comprise a recombinase complex associated with the MG64 Cas effector.

[0050] SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222 show the nucleotide sequences of MG64 tracrRNA derived from the same locus as the MG64 Cas effector.

[0051] SEQ ID NOs: 7 and 34-35 show the nucleotide sequences of MG64-targeted CRISPR repeats.

[0052] SEQ ID NOs: 106-108, 112-118, and 221 show the nucleotide sequences of MG64 crRNA.

[0053] SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93 show the nucleotide sequences of the right transporase recognition sequences associated with the MG64 system.

[0054] SEQ ID NOs: 9, 11, 36-38, 76, and 78 show the nucleotide sequences of the left-side transposase recognition sequences associated with the MG64 system.

[0055] SEQ ID NO:31 shows the PAM sequence associated with the MG64 Cas effector described herein.

[0056] SEQ ID NOs: 45-63, 68-75, 96-103, and 123-140 show the nucleotide sequences of single guide RNAs engineered to function with the MG64 Cas effector.

[0057] SEQ ID NO: 208 shows the nucleotide sequence of the MG64 expression construct.

[0058] MG190 SEQ ID NOs: 209 to 219 show the full-length peptide sequences of MG190 ribosomal protein S15 homologues.

[0059] Other Arrays SEQ ID NOs: 86 to 87 and 192 to 207 show peptide sequences of nuclear localization signals.

[0060] SEQ ID NOs: 88 to 89 show the peptide sequences of linkers.

[0061] SEQ ID NOs: 90 to 92 show the peptide sequences of epitope tags.

[0062] SEQ ID NOs: 141 to 143 show genomic target sequences.

[0063] SEQ ID NOs: 144 to 180 show targeting guide sequences.

[0064] SEQ ID NOs: 181 to 183 show the nucleic acid sequences of S15 fusion proteins.

[0065] SEQ ID NO: 184 shows the donor construct.

[0066] SEQ ID NO: 185 shows the MG64-1 sgRNA sequence.

[0067] SEQ ID NO: 186 shows the linker sequence.

[0068] SEQ ID NOs: 187 to 189 show the amino acid sequences of S15 fusion proteins.

[0069] SEQ ID NOs: 190 to 191 show promoter sequences. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0070] While various embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the present disclosure. It should be understood that various alternatives to the embodiments of the present disclosure described herein may be used.

[0071] The practice of some of the methods disclosed herein, unless otherwise indicated, employs techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA.See, for example, Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012), the series Current Protocols in Molecular Biology (FM Ausubel, et al. eds.), the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (MJ MacPherson, BD Hames and GR Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (RI Freshney, ed. (2010)).

[0072] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. Furthermore, to the extent the terms "including," "includes," "having," "has," "with," or variations thereof are used in any of the detailed description and / or claims, such terms are intended to be inclusive in the same manner as the term "comprising."

[0073] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within one or more standard deviations, as is customary in the art. Alternatively, "about" can mean within a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of a given value.

[0074] As used herein, "cell" refers to a biological cell. A cell may be the basic structural, functional, and / or biological unit of a living organism. A cell may originate from any organism having one or more cells. Some non-limiting examples include prokaryotic cells, eukaryotic cells, bacterial cells, archaeal cells, single-cell eukaryotic cells, protozoan cells, cells from plants (e.g., plant crops, fruits, vegetables, grains, soybeans, cane, corn, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkins, hay, potatoes, cotton, cannabis, tobacco, flowering plants, conifers, gymnosperms, ferns, club mosses, hornworts, mosses, cells from mosses), algae cells (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, etc.), and cells from plants (e.g., cereals, fruits, vegetables, grains, soybeans, cane, corn, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkins, hay, potatoes, cotton, cannabis, tobacco, flowering plants, conifers, gymnosperms, ferns, club mosses, hornworts, mosses, cells from mosses), algae cells (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, etc.). C. Agardh, etc.), seaweed (e.g., kelp), fungal cells (e.g., yeast cells, cells from mushrooms), animal cells, cells from invertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), cells from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), cells from mammals (e.g., pigs, cows, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.). Sometimes the cells are not derived from a naturally occurring organism (e.g., cells can be synthetically produced and sometimes referred to as artificial cells).

[0075] The term "nucleotide" as used herein refers to a base-sugar-phosphate combination. Nucleotides may include synthetic nucleotides. Nucleotides may include synthetic nucleotide analogs. Nucleotides may be monomeric units of nucleic acid sequences (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide may include ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP) and deoxyribonucleoside triphosphates, such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives may include, for example, [αS]dATP, 7-deaza-dGTP and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules that contain them. The term nucleotide as used herein may refer to dideoxyribonucleoside triphosphates (ddNTPs) and derivatives thereof. Illustrative examples of dideoxyribonucleoside triphosphates may include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides may be unlabeled or detectably labeled, such as by using a moiety that includes an optically detectable moiety (e.g., a fluorophore). Labeling may also be performed using quantum dots. Detectable labels may include, for example, radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels. Fluorescent labels for nucleotides may include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2′7′-dimethoxy-4′5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N′,N′-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4′dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2′-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP available from Perkin Elmer (Foster City, Calif.); fluoro-conjugated deoxynucleotides, fluoro-conjugated Cy3-dCTP, fluoro-conjugated Cy5-dCTP, fluoro-conjugated fluoroX-dCTP, fluoro-conjugated Cy3-dUTP, and fluoro-conjugated Cy5-dUTP available from Amersham (Arlington Heights, Ill.); Fluorescein-15-dATP, fluorescein-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2′-dATP available from Mannheim (Indianapolis, Ind.); and Molecular Examples of chromosomal labeling nucleotides available from Probes (Eugene, Oreg.) include BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP. Nucleotides can also be labeled or marked by chemical modification. The chemically modified single nucleotide may be a biotin-dNTP.Some non-limiting examples of biotinylated dNTPs can include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).

[0076] The terms "polynucleotide", "oligonucleotide", and "nucleic acid" are used interchangeably to refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof, in single-stranded, double-stranded, or multiple-stranded form. A polynucleotide may be exogenous or endogenous to a cell. A polynucleotide may be present in a cell-free environment. A polynucleotide may be a gene or a fragment thereof. A polynucleotide may be DNA. A polynucleotide may be RNA. A polynucleotide may have any three-dimensional structure and perform any function. In polynucleotides when referring to T, T means U (uracil) in RNA and T (thymine) in DNA. A polynucleotide may contain one or more analogs (e.g., modified backbones, sugars, or nucleobases). If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. Some non-limiting examples of analogs include 5-bromouracil, peptide nucleic acid, heterologous nucleic acid, morpholino, locked nucleic acid, glycol nucleic acid, threose nucleic acid, dideoxynucleotides, cordycepin, 7-diaza-GTP, fluorophores (e.g., rhodamine or fluorescein attached to the sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queosine, and wyosine. Non-limiting examples of polynucleotides include coding or non-coding regions of a gene or gene fragment, loci defined from binding analyses, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides, including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers.The sequence of nucleotides may be interrupted by non-nucleotide components.

[0077] The term "transfection" or "transfected" refers to the introduction of a nucleic acid into a cell by non-viral or viral-based methods. The nucleic acid molecule can be a genetic sequence encoding a complete protein or a functional portion thereof. See, e.g., Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, 18.1-18.88.

[0078] The terms "peptide", "polypeptide", and "protein" are used interchangeably herein to refer to a polymer of at least two amino acid residues linked by peptide bonds. The term does not refer to a specific length of the polymer, nor is it intended to imply or distinguish whether the peptide is produced using recombinant technology, chemical or enzymatic synthesis, or naturally occurring. The term applies to naturally occurring amino acid polymers as well as amino acid polymers that contain at least one modified amino acid. In some cases, the polymer may be interrupted by non-amino acids. The term includes amino acid chains of any length, including full-length proteins, and proteins with or without secondary and / or tertiary structure (e.g., domains). The term also encompasses amino acid polymers that have been modified by any other manipulation, such as, for example, disulfide bond formation, glycosylation, lipid formation, acetylation, phosphorylation, oxidation, and conjugation with a labeling component. The terms "amino acid" and "amino acids" as used herein refer to natural and unnatural amino acids, including, but not limited to, modified amino acids and amino acid analogs. Modified amino acids can include natural and unnatural amino acids that have been chemically modified to include non-naturally occurring groups or chemical moieties on the amino acid. Amino acid analogs can refer to amino acid derivatives. The term "amino acid" includes both D- and L-amino acids.

[0079] As used herein, "non-natural" may refer to a nucleic acid or polypeptide sequence that is not found in a natural nucleic acid or protein. Non-natural may refer to an affinity tag. Non-natural may refer to a fusion. Non-natural may refer to a naturally occurring nucleic acid or polypeptide sequence that includes mutations, insertions, and / or deletions. A non-natural sequence may exhibit and / or encode an activity (e.g., an enzyme activity, a methyltransferase activity, an acetyltransferase activity, a kinase activity, an ubiquitination activity, etc.) that may also be exhibited by the nucleic acid and / or polypeptide sequence to which the non-natural sequence is fused. A non-natural nucleic acid or polypeptide sequence may be joined by genetic engineering to a naturally occurring nucleic acid and / or polypeptide sequence (or a variant thereof) to generate a chimeric nucleic acid and / or polypeptide sequence that encodes a chimeric nucleic acid or polypeptide.

[0080] The term "promoter" as used herein refers to a regulatory DNA region that controls the transcription or expression of a polynucleotide (e.g., a gene) and may be located adjacent to or overlapping the nucleotide or region of nucleotides at which RNA transcription is initiated. A promoter may contain specific DNA sequences that bind protein factors, often referred to as transcription factors, that facilitate the binding of RNA polymerase to DNA resulting in gene transcription. A "basal promoter", also referred to as a "core promoter", may refer to a promoter that contains all the basic and necessary elements to facilitate the transcriptional expression of an operably linked polynucleotide. Eukaryotic basal promoters typically, but not necessarily, contain a TATA-box and / or a CAAT box. In some embodiments, different promoters induce the expression of a gene in different tissues or cell types, or at different developmental stages, or in response to different environmental or physiological conditions or inducer molecules. A promoter that causes a gene to be expressed in most cell types in most cases is commonly referred to as a constitutive promoter. A promoter that causes a gene to be expressed in specific cell and tissue types is commonly referred to as a "cell-specific promoter" or a "tissue-specific promoter", respectively. Promoters that cause expression of a gene at a specific stage of development or cell differentiation are generally referred to as "development-specific promoters" or "cell differentiation-specific promoters." Promoters that induce and result in expression of a gene after exposure or treatment of cells with a promoter-inducing drug, biomolecule, chemical, ligand, light, etc. are generally referred to as "inducible promoters" or "regulatable promoters." It is further recognized that in some embodiments, DNA fragments of different lengths have the same promoter activity, since the exact boundaries of regulatory sequences are in most cases not completely defined.

[0081] The term "expression" as used herein refers to the process by which a nucleic acid sequence or polynucleotide is transcribed from a DNA template (e.g., into mRNA or other RNA transcript) and / or the process by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide may be collectively referred to as the "gene product." If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell.

[0082] As used herein, "operably linked", "operable linkage", "operatively linked", or their grammatical equivalents refer to an arrangement of genetic elements, such as promoters, enhancers, polyadenylation sequences, etc., where the action (e.g., movement or activation) of a first genetic element has some effect on a second genetic element. The effect on the second genetic element can be, but need not be, of the same type as the action of the first genetic element. For example, two genetic elements are operably linked if movement of the first element causes activation of the second element. A regulatory element is operably linked to a coding region if the regulatory element helps to initiate transcription of the coding sequence, which may include, for example, a promoter sequence and / or an enhancer sequence. There may be intervening residues between the regulatory element and the coding region, so long as this functional relationship is maintained.

[0083] As used herein, a "vector" refers to a polymer or association of polymers that contains or is associated with a polynucleotide and can be used to mediate delivery of the polynucleotide to a cell. Examples of vectors include plasmids, viral vectors, liposomes, and other gene delivery vehicles. A vector generally includes a genetic element, e.g., a regulatory element, operably linked to a gene to facilitate expression of the gene in a target.

[0084] As used herein, "expression cassette" and "nucleic acid cassette" are used interchangeably to refer to a combination of nucleic acid sequences or elements that are expressed together or operably linked for expression. In some cases, an expression cassette refers to a combination of regulatory elements and one or more genes that are operably linked for expression.

[0085] A "functional fragment" of a DNA or protein sequence refers to a fragment that retains a biological activity (either functional or structural) substantially similar to that of the full-length DNA or protein sequence. The biological activity of a DNA sequence can be the ability to affect expression in a manner attributable to the full-length sequence.

[0086] The terms "engineered," "synthetic," and "artificial" are used interchangeably herein to refer to an entity that has been modified by human intervention. For example, the terms may refer to a polynucleotide or polypeptide that does not occur in nature. An engineered peptide may, but need not, have low sequence identity (e.g., less than 50% sequence identity, less than 25% sequence identity, less than 10% sequence identity, less than 5% sequence identity, less than 1% sequence identity) with a naturally occurring human protein. For example, the VPR domain and the VP64 domain are synthetic transactivation domains. By way of non-limiting examples, a nucleic acid may be modified by changing its sequence to a sequence that does not occur in nature, a nucleic acid may be modified by ligating to a nucleic acid that is not naturally associated with it such that the ligated product possesses a function not present in the original nucleic acid, an engineered nucleic acid may be synthesized in vitro with a sequence that does not occur in nature, a protein may be modified by changing its amino acid sequence to a sequence that does not occur in nature, an engineered protein may acquire a new function or property. An "engineered" system includes at least one engineered component.

[0087] The term "tracrRNA" or "tracr sequence" refers to transactivating CRISPR RNA. tracrRNA interacts with CRISPR(cr)RNA to form a guide nucleic acid (e.g., guide RNA or gRNA) that can hybridize to a target nucleic acid and thereby direct an associated nuclease to the target nucleic acid. When tracrRNA is engineered, it can have about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% sequence identity and / or sequence similarity with a wild-type exemplary tracrRNA sequence (e.g., tracrRNA from S. pyogenes, S. aureus, etc.). tracrRNA can refer to modified forms of tracrRNA that can include nucleotide changes such as deletions, insertions, or substitutions, variants, mutations, or chimeras. tracrRNA may refer to a nucleic acid that may be at least about 60% identical to a wild-type exemplary tracrRNA (e.g., tracrRNA from S. pyogenes, S. aureus, etc.) sequence over a stretch of at least six consecutive nucleotides. For example, a tracrRNA sequence may be at least about 60% identical, at least about 65% identical, at least about 70% identical, at least about 75% identical, at least about 80% identical, at least about 85% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, or 100% identical to a wild-type exemplary tracrRNA (e.g., tracrRNA from S. pyogenes, S. aureus, etc.) sequence over a stretch of at least six consecutive nucleotides. Type II tracrRNA sequences can be predicted on genomic sequences by identifying regions that have complementarity to portions of repeat sequences in adjacent CRISPR arrays.

[0088] As used herein, a "guide nucleic acid" or "guide polynucleotide" refers to a nucleic acid that can hybridize to a target nucleic acid and thereby direct an associated nuclease to the target nucleic acid. A guide nucleic acid can be an RNA (guide RNA or gRNA). A guide nucleic acid can be a DNA. A guide nucleic acid can be a mixture of RNA and DNA. A guide nucleic acid can include crRNA or tracrRNA, or a combination of both. A guide nucleic acid can be engineered. A guide nucleic acid can be programmed to specifically bind to a target nucleic acid. A portion of a target nucleic acid can be complementary to a portion of a guide nucleic acid. A strand of a double-stranded target polynucleotide that is complementary to a guide nucleic acid and hybridizes with the guide nucleic acid can be referred to as a complementary strand. A strand of a double-stranded target polynucleotide that is complementary to a complementary strand and therefore may not be complementary to the guide nucleic acid can be referred to as a non-complementary strand. A guide nucleic acid can include a polynucleotide strand and can be referred to as a "single guide nucleic acid." A guide nucleic acid can include two polynucleotide strands and can be referred to as a "double guide nucleic acid." Unless otherwise specified, the term "guide nucleic acid" is inclusive and can refer to both single and double guide nucleic acids. A guide nucleic acid can include a segment that can be referred to as a "nucleic acid targeting segment" or a "nucleic acid targeting sequence" or a "spacer sequence." A nucleic acid targeting segment can include a sub-segment that can be referred to as a "protein binding segment" or a "protein binding sequence" or a "Cas protein binding segment."

[0089] As used herein, the terms "gene editing" and "genome editing" can be used interchangeably. Gene editing or genome editing refers to changing the nucleic acid sequence of a gene or genome. Genome editing can include, for example, insertions, deletions, and mutations.

[0090] The term "sequence identity" or "percent identity" in the context of two or more nucleic acid or polypeptide sequences refers to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences that are identical, or have a certain percentage of identical amino acid residues or nucleotides, when compared and aligned for maximum correspondence over a local or global comparison window, as measured using a sequence comparison algorithm. Suitable sequence comparison algorithms for polypeptide sequences include, for example, BLASTP using the BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and presence of 11, gap cost at an extension of 1, and using a conditional composition score matrix adjustment for polypeptide sequences longer than 30 residues; BLASTP using parameters of word length (W) of 2, expectation (E) of 1,000,000, and PAM30 scoring setting gap costs at 9 for open gaps and 1 for extended gaps for sequences shorter than 30 residues (default parameters for BLASTP are available in BLAST at https: / / blast.ncbi.nlm.nih.gov); CLUSTALW using the Smith-Waterman homology search algorithm with parameters of match of 2, mismatch of -1, and gap of -1; MUSCLE with default parameters; MAFFT with parameters retree of 2 and maximum iterations of 1000; Novafold with default parameters; HMMER hmmalign with default parameters.

[0091] The present disclosure includes variants of any of the enzymes described herein that have one or more conservative amino acid substitutions. Such conservative substitutions can be made in the amino acid sequence of a polypeptide without destroying the three-dimensional structure or function of the polypeptide. Conservative substitutions can be achieved by substituting amino acids with similar hydrophobicity, polarity, and R chain length for each other. Additionally or alternatively, by comparing aligned sequences of homologous proteins from different species, conservative substitutions can be identified by identifying amino acid residues that are not mutated between species (e.g., residues that are not conserved without altering the basic function of the encoded protein). Such conservatively substituted variants may include variants having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of the systems described herein (e.g., the MG64 system described herein). In some embodiments, such conservatively substituted variants are functional variants. Such functional variants may include sequences with substitutions such that the activity of key active site residues of the endonuclease is not destroyed. In some embodiments, a functional variant of any of the systems described herein lacks at least one substitution of a conserved or functional residue called out in Figures 4A, 4B, and 5. In some embodiments, a functional variant of any of the systems described herein lacks all substitutions of a conserved or functional residue called out in Figures 4A, 4B, and 5.

[0092] Conservative substitution tables providing functionally similar amino acids are available in a variety of references (e.g., Creighton, Proteins: Structures and Molecular Properties (WH Freeman & Co.; 2 nd The following eight groups each contain amino acids that are conservative substitutions for one another: 1) Alanine (A), Glycine (G); 2) Aspartic acid (D), glutamic acid (E); 3) Asparagine (N), Glutamine (Q); 4) arginine (R), lysine (K); 5) isoleucine (I), leucine (L), methionine (M), valine (V); 6) phenylalanine (F), tyrosine (Y), tryptophan (W); 7) serine (S), threonine (T); and 8) Cysteine ​​(C), Methionine (M).

[0093] As used herein, the term "RuvC_III domain" refers to the third discontinuous segment of the RuvC endonuclease domain (the RuvC nuclease domain is composed of three discontinuous segments, RuvC_I, RuvC_II, and RuvC_III). RuvC domains or segments thereof can generally be identified by alignment to documented domain sequences, structural alignment to proteins with annotated domains, or comparison to hidden Markov models (HMMs) built based on documented domain sequences (e.g., Pfam HMM PF18541 for RuvC_III).

[0094] As used herein, the term "HNH domain" refers to an endonuclease domain having characteristic histidine and asparagine residues. HNH domains can generally be identified by alignment to documented domain sequences, structural alignment to proteins with annotated domains, or comparison to hidden Markov models (HMMs) constructed based on documented domain sequences (e.g., Pfam HMM PF01844 for domain HNH).

[0095] As used herein, the term "recombinase" refers to an enzyme that mediates the recombination of DNA fragments located between recombinase recognition sequences, resulting in excision, insertion, inversion, exchange, or transposition) of the DNA fragments located between the recombinase recognition sequences.

[0096] As used herein, the term "recombining" or "recombination" in the context of nucleic acid modification (e.g., genomic modification) refers to a process in which two or more nucleic acid molecules, or two or more regions of a single nucleic acid molecule, are modified by the action of a recombinase protein. Recombination can result in, among other things, excision, insertion, inversion, exchange, or rearrangement of a nucleic acid sequence within or between one or more nucleic acid molecules.

[0097] As used herein, the term "transposon" or "transposable element" refers to a nucleic acid sequence in a genome that is a mobile genetic element that can change its position in the genome. In some cases, the transposon transports additional "cargo DNA" that is excised from the genome. Transposons include, for example, retrotransposons, DNA transposons, autonomous and non-autonomous transposons, and class III transposons. The transposon nucleic acid sequence includes, for example, a gene encoding a cognate transposase, one or more recognition sequences for the transposase, or a combination thereof. In some cases, these transposons differ by the type of nucleic acid that they transpose, the type of repeats at the ends of the transposon, the type of cargo they carry, or the mode of transposition (i.e., self-repair or host repair). As used herein, the term "transposase" or "transposases" refers to an enzyme that binds to the recognition sequence of the transposon and catalyzes its movement to another part of the genome. In some cases, the movement is by a cut-and-paste mechanism or a copy-and-transposition mechanism.

[0098] As used herein, the term "Tn7" or "Tn7-like transposase" refers to a family of transposases that includes three major components: a heteromeric transposase (TnsA and / or TnsB) along with a regulatory protein (TnsC). In addition to the TnsABC transposition proteins, Tn7 elements can encode dedicated target site selection proteins, TnsD and TnsE. In conjunction with TnsABC, the sequence-specific DNA binding protein TnsD directs transposition into a conserved site called the "Tn7 attachment site", i.e., attTn7. TnsD is a member of a large family of proteins that also includes TniQ. TniQ has been shown to target transposition into the degradation site of a plasmid.

[0099] As used herein, the term "complex" refers to the conjugation of at least two components. Each of the two components may retain the properties / activity it had before forming the complex. The conjugation may be by covalent bonds, non-covalent bonds (i.e., hydrogen bonds, ionic interactions, van der Waals interactions, and hydrophobic bonds), use of linkers, fusion, or any other suitable method. In some cases, the components in the complex are polynucleotides, polypeptides, or combinations thereof. For example, the complex may include a Cas protein and a guide nucleic acid.

[0100] In some cases, the CAST system described herein comprises one or more Tn7 or Tn7-like transposases. In certain exemplary embodiments, the Tn7 or Tn7-like transposase comprises a multimeric protein complex. In certain exemplary embodiments, the multimeric protein complex comprises TnsA, TnsB, TnsC, or TniQ. In these combinations, the transposases (TnsA, TnsB, TnsC, TniQ) may form a complex or fusion protein with each other.

[0101] In some cases, the CAST system described herein comprises one or more Tn5053 or Tn5053-like transposases. In certain exemplary embodiments, the Tn5053 or Tn5053-like transposase comprises a multimeric protein complex. In certain exemplary embodiments, the multimeric protein complex comprises TnsA, TnsB, TnsC, or TniQ. In these combinations, the transposases (TnsA, TnsB, TnsC, TniQ) may form a complex or fusion protein with each other.

[0102] As used herein, the term "Cas12k" (alternatively "Class 2 VK type") refers to a subtype of V type CRISPR system that has been found to be defective in nuclease activity (e.g., they may contain at least one defective RuvC domain that lacks at least one catalytic residue important for DNA cleavage). Such effector subtypes are generally associated with the CAST system.

[0103] In accordance with IUPAC convention, the following abbreviations are used throughout the examples: A=Adenine C=Cytosine G=guanine T=Thymine R = adenine or guanine Y = cytosine or thymine S = guanine or cytosine W = adenine or thymine K = guanine or thymine M = adenine or cytosine B=C, G, or T D=A, G, or T H=A, C, or T V=A, C, or G

[0104] overview Discovery of new Cas enzymes with unique functionality and structure could confer the potential to further disrupt deoxyribonucleic acid (DNA) editing technologies, improving their speed, specificity, functionality, and ease of use. Compared to the predicted prevalence of clustered regularly interspaced short palindromic repeats (CRISPR) systems in microbes and the net diversity of microbial species, a relatively small number of functionally characterized CRISPR / Cas enzymes exist in the literature. This is in part because the vast number of microbial species are not easily cultured under laboratory conditions. Metagenomic sequencing from natural environmental niches representing a large number of microbial species could dramatically increase the number of documented new CRISPR / Cas systems and confer the potential to expedite the discovery of new oligonucleotide editing functions. A fruitful recent example of such an approach is demonstrated by the 2016 discovery of the CasX / CasY CRISPR system from metagenomic analysis of natural microbial communities.

[0105] CRISPR / Cas systems are RNA-directed nuclease complexes that have been described to function as adaptive immune systems in microorganisms. In their natural context, CRISPR / Cas systems occur in CRISPR (clustered regularly interspaced short palindromic repeats) operons or loci, which generally contain two parts: (i) an array of short repetitive sequences (30-40 bp) separated by equally short spacer sequences that encode RNA-based targeting elements, and (ii) an ORF encoding a Cas that encodes a nuclease polypeptide that is directed by the RNA-based targeting element flanked by accessory proteins / enzymes. Efficient nuclease targeting of a specific target nucleic acid sequence generally requires both (i) complementary hybridization between the first 6-8 nucleic acids of the target (target seed) and the crRNA guide; and (ii) the presence of a protospacer adjacent motif (PAM) sequence within a defined vicinity of the target seed (PAM is usually a sequence that is not commonly represented in the host genome). Depending on the exact function and composition of the system, CRISPR-Cas systems are commonly organized into two classes, five types, and 16 subtypes based on shared functional characteristics and evolutionary similarities (see Figure 1).

[0106] Class 1 CRISPR-Cas systems have large multi-subunit effector complexes and include types I, III, and IV.

[0107] Type I CRISPR-Cas systems are considered to be of intermediate complexity in terms of components. In type I CRISPR-Cas systems, an array of RNA targeting elements is transcribed as a long precursor crRNA (pre-crRNA) that is processed at the repeat elements to release a short mature crRNA that directs the nuclease complex to the nucleic acid target, followed by a suitable short consensus sequence called the protospacer adjacent motif (PAM). This processing occurs via the endoribonuclease subunit (Cas6) of a large endonuclease complex called Cascade, which also contains the nuclease (Cas3) protein component of the crRNA-directed nuclease complex. Cas I nuclease functions primarily as a DNA nuclease.

[0108] Type III CRISPR systems can be characterized by the presence of a central nuclease known as Cas10, along with repeat-associated mysterious proteins (RAMPs) that contain Csm or Cmr protein subunits. Similar to type I systems, mature crRNA is processed from pre-crRNA using a Cas6-like enzyme. Unlike type I and type II systems, type III systems appear to target and cleave DNA-RNA duplexes (such as the DNA strand used as a template for RNA polymerase).

[0109] Type IV CRISPR-Cas systems possess an effector complex that contains a highly reduced large subunit nuclease (csf1), two genes for RAMP proteins of the Cas5 (csf3) and Cas7 (csf2) family, and in some cases a predicted small subunit gene; such systems are commonly found on endogenous plasmids.

[0110] Class 2 CRISPR-Cas systems generally have a single polypeptide multi-domain nuclease effector and include Types II, V, and VI.

[0111] Type II CRISPR-Cas systems are considered the simplest in terms of components. In type II CRISPR-Cas systems, the processing of the CRISPR array into mature crRNA does not require the presence of a special endonuclease subunit, but rather a small transcoding crRNA (tracrRNA) with a region complementary to the array repeat sequence, which interacts with both its corresponding effector nuclease (e.g., Cas9) and the repeat sequence to form a precursor dsRNA structure that is cleaved by endogenous RNAse III to generate the mature effector enzyme loaded with both tracrRNA and crRNA. Type II nucleases are known as DNA nucleases. Type II effectors generally exhibit a structure consisting of a RuvC-like endonuclease domain that adopts an RNase H fold with an unrelated HNH nuclease domain inserted into the fold of the RuvC-like nuclease domain. The RuvC-like domain is involved in cleavage of the target (e.g., crRNA-complementary) DNA strand, while the HNH domain is involved in cleavage of the variant DNA strand.

[0112] Type V CRISPR-Cas systems are characterized by a nuclease effector (e.g., Cas12) structure similar to that of type II effectors, including a RuvC-like domain. Like type II, most (if not all) type V CRISPR systems use tracrRNA to process pre-crRNA into mature crRNA, but unlike type II systems that require RNAse III to cleave pre-crRNA into multiple crRNAs, type V systems are capable of cleaving pre-crRNA using the effector nuclease itself. Like type II CRISPR-Cas systems, type V CRISPR-Cas systems are again known as DNA nucleases. Unlike type II CRISPR-Cas systems, some type V enzymes (e.g., Cas12a) appear to have robust single-stranded non-specific deoxyribonuclease activity that is activated by the first crRNA-directed cleavage of the double-stranded target sequence.

[0113] Type VI CRISPR-Cas systems have an RNA-guided RNA endonuclease. Instead of a RuvC-like domain, the single polypeptide effector of type VI systems (e.g., Cas13) contains two HEPN ribonuclease domains. Unlike both type II and type V systems, type VI systems also do not appear to require tracrRNA to process pre-crRNA into crRNA. However, similar to type V systems, some type VI systems (e.g., C2C2) appear to have robust single-stranded non-specific nuclease (ribonuclease) activity that is activated by the first crRNA-directed cleavage of the target RNA.

[0114] Due to their simpler structure, class 2 CRISPR-Cas have been most widely adopted for engineering and development as engineered nuclease / genome editing applications.

[0115] One of the initial adaptations of such a system for in vivo use involves the use of (i) purified recombinantly expressed full-length Cas9 (e.g., a class 2 type II Cas enzyme) isolated from S. pyogenes SF370, (ii) purified mature, approximately 42 nt crRNA (total crRNA transcribed in vitro from a synthetic DNA template carrying a T7 promoter sequence) carrying an approximately 20 nt 5' sequence complementary to the target DNA sequence desired to be cleaved, followed by a 3' tracr binding sequence, (iii) purified tracrRNA transcribed in vitro from a synthetic DNA template carrying a T7 promoter sequence, and (iv) Mg 2+ Subsequent improved and engineered systems involved a crRNA (ii) joined to the 5' end of (iii) by a linker (e.g., GAAA) to form a single fusion synthetic guide RNA (sgRNA) that can itself guide Cas9 to the target (compare the top and bottom panels of FIG. 2).

[0116] Such engineered systems can be adapted for use in mammalian cells by providing a DNA vector encoding (i) an ORF encoding a codon-optimized Cas9 (e.g., a class 2 type II Cas enzyme) under a suitable mammalian promoter with a C-terminal nuclear localization sequence (e.g., SV40 NLS) and a suitable polyadenylation signal (e.g., TK pA signal), and (ii) an ORF encoding an sgRNA (having a 5' sequence starting with G, followed by 20 nt of complementary targeting nucleic acid sequence joined to the 3' tracr binding sequence, a linker, and the tracrRNA sequence) under a suitable polymerase III promoter (e.g., U6 promoter).

[0117] Transposons are mobile elements that can move between locations in a genome. Such transposons have evolved to limit the negative effects they have on the host. Various control mechanisms are used to maintain translocation at low frequency and sometimes to coordinate translocation with various cellular processes. Some prokaryotic transposons can also marshal functions that benefit the host or otherwise help maintain the element. Certain transposons may also have evolved mechanisms of strict control over target site selection, the most prominent example being the Tn7 family.

[0118] Transposon Tn7 and similar elements are reservoirs for antibiotic resistance and pathogenic functions in clinical settings, and may encode other adaptive functions in the natural environment. The Tn7 system, for example, has evolved mechanisms to almost completely avoid integrating into critical host genes, but to maximize dispersal of the element by recognizing mobile plasmids and bacteriophages capable of transferring Tn7 between host bacteria.

[0119] Tn7 and Tn7-like elements control where and when they insert, and may possess one pathway that directs insertion into a single conserved location in the bacterial genome, and a second pathway that appears to be adapted to maximize targeting into mobile plasmids capable of transporting elements between bacteria (Figure 3). The link between Tn7-like transposons and CRISPR-Cas systems suggests that transposons may have hijacked CRISPR effectors that generate R-loops at target sites, facilitating the spread of transposons through plasmids and phages.

[0120] MG64 series In some embodiments, provided herein is an MG64 system for transposing a cargo nucleotide sequence into a target nucleic acid site. See Figures 4A-4B. In some embodiments, the system comprises a double-stranded nucleic acid comprising a cargo nucleotide sequence. In some embodiments, the cargo nucleotide sequence is configured to interact with a Tn7-type or Tn5053-type transposase complex. In some embodiments, the system comprises a Cas effector complex. In some embodiments, the Cas effector complex comprises a class 2 V-type Cas effector and an engineered guide polynucleotide configured to hybridize to a target nucleotide sequence. In some embodiments, the system comprises a Tn7-type or Tn5053-type transposase complex configured to bind to the Cas effector complex, the Tn7-type or Tn5053-type transposase complex comprising a TnsB subunit.

[0121] In some cases, the cargo nucleotide sequence is adjacent to a left transposase recognition sequence. In some cases, the cargo nucleotide sequence is adjacent to a right transposase recognition sequence. In some cases, the cargo nucleotide sequence is adjacent to a left transposase recognition sequence and a right transposase recognition sequence.

[0122] In some cases, the target nucleic acid comprises a target nucleic acid site. In some cases, the target nucleic acid comprises a PAM sequence that is compatible with a Cas effector complex adjacent to the target nucleic acid site. In some cases, the PAM sequence is located 3' of the target nucleic acid site. In some cases, the PAM sequence is located 5' of the target nucleic acid site.

[0123] In some cases, the engineered guide polynucleotide is configured to bind to a class 2 V-type Cas endonuclease. In some cases, the class 2 V-type Cas effector is a class 2 VK-type effector. In some cases, the Class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NOs:1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 70% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 75% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 80% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 85% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least about 90% identity to SEQ ID NO:1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least about 91% identity to SEQ ID NO:1, 12, 16, 20-30, 64, 80-85, and 220.In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 92% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 93% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 94% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 95% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 96% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 97% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 98% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 99% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having 100% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220.

[0124] In some cases, the TnsB subunit comprises a polypeptide having a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB subunit comprises a polypeptide having a sequence identical to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 70% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 75% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 80% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 85% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 90% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 91% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 92% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 93% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 94% identity to SEQ ID NOs: 2, 13, 17, and 65.In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 95% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 96% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 97% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 98% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 99% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having 100% identity to SEQ ID NOs: 2, 13, 17, and 65.

[0125] In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs:3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 70% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 75% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 80% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 85% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 90% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 91% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 92% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 93% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 94% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 95% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 96% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 97% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 98% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 99% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having 100% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.

[0126] In some cases, the Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least one of SEQ ID NOs:3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least 70% sequence identity to at least one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently having at least about 85% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least about 90% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least about 91% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least about 92% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least about 93% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least about 94% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least about 95% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least about 96% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least about 97% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least about 98% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least about 99% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least 100% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.

[0127] In some embodiments, the systems disclosed herein comprise at least one engineered guide polynucleotide, e.g., a gRNA.

[0128] In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 70% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 75% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 80% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 85% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 90% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222.In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 91% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 92% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 93% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 94% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 95% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 96% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 97% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 98% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222.In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 99% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having 100% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222.

[0129] In some cases, the engineered guide RNA comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, 123-140, and 185. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides identical to any one of SEQ ID NOs: 45-63, 68-75, 96-103, 123-140, and 185. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 70% identity to SEQ ID NOs: 45-63, 68-75, 96-103, 123-140, and 185. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 75% identity to SEQ ID NOs: 45-63, 68-75, 96-103, 123-140, and 185. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 80% identity to SEQ ID NOs: 45-63, 68-75, 96-103, 123-140, and 185. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 85% identity to SEQ ID NOs: 45-63, 68-75, 96-103, 123-140, and 185.In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that comprises at least about 46-80 contiguous nucleotides that have at least about 90% identity to SEQ ID NOs: 45-63, 68-75, 96-103, 123-140, and 185. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that comprises at least about 46-80 contiguous nucleotides that have at least about 91% identity to SEQ ID NOs: 45-63, 68-75, 96-103, 123-140, and 185. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that comprises at least about 46-80 contiguous nucleotides that have at least about 92% identity to SEQ ID NOs: 45-63, 68-75, 96-103, 123-140, and 185. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that comprises at least about 46-80 contiguous nucleotides that have at least about 93% identity to SEQ ID NOs: 45-63, 68-75, 96-103, 123-140, and 185. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that comprises at least about 46-80 contiguous nucleotides that have at least about 94% identity to SEQ ID NOs: 45-63, 68-75, 96-103, 123-140, and 185. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that comprises at least about 46-80 contiguous nucleotides that have at least about 95% identity to SEQ ID NOs: 45-63, 68-75, 96-103, 123-140, and 185. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 96% identity to SEQ ID NOs: 45-63, 68-75, 96-103, 123-140, and 185. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 97% identity to SEQ ID NOs: 45-63, 68-75, 96-103, 123-140, and 185.In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that comprises at least about 46-80 contiguous nucleotides that have at least about 98% identity to SEQ ID NOs: 45-63, 68-75, 96-103, 123-140, and 185. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that comprises at least about 46-80 contiguous nucleotides that have at least about 99% identity to SEQ ID NOs: 45-63, 68-75, 96-103, 123-140, and 185. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that comprises at least about 46-80 contiguous nucleotides that have 100% identity to SEQ ID NOs: 45-63, 68-75, 96-103, 123-140, and 185.

[0130] In some embodiments, the guide RNA comprises various structural elements, including but not limited to a spacer sequence that binds to a protospacer sequence (target sequence), a crRNA, and an optional tracrRNA. In some embodiments, the guide RNA comprises a crRNA that comprises a spacer sequence. In some embodiments, the guide RNA additionally comprises a tracrRNA or a modified tracrRNA.

[0131] In some embodiments, the systems provided herein include one or more guide RNAs. In some embodiments, the guide RNA includes a sense sequence. In some embodiments, the guide RNA includes an antisense sequence. In some embodiments, the guide RNA includes a nucleotide sequence other than a region that is complementary or substantially complementary to a region of a target sequence. For example, the crRNA is or is considered to be part of the guide RNA, or is included in the guide RNA, e.g., a crRNA:tracrRNA chimera.

[0132] In some embodiments, the guide RNA comprises synthetic or modified nucleotides. In some embodiments, the guide RNA comprises one or more internucleoside linkers modified from natural phosphodiester. In some embodiments, the internucleoside linker of the guide RNA, or all of its contiguous nucleotide sequence, is modified. For example, in some embodiments, the internucleoside linkage comprises sulfur (S), such as a phosphorothioate internucleoside linkage.

[0133] In some embodiments, the guide RNA comprises a modification to the ribose sugar or nucleobase. In some embodiments, the guide RNA comprises one or more nucleosides comprising a modified sugar moiety, which is a modification of the sugar moiety as compared to the ribose sugar moiety found in deoxyribose nucleic acids (DNA) and RNA. In some embodiments, the modification is in the ribose ring structure. Exemplary modifications include, but are not limited to, replacement with a hexose ring (HNA), a bicyclic ring having a biradical bridge between the C2 and C4 carbons on the ribose ring (e.g., locked nucleic acid (LNA)), or a non-linked ribose ring that typically lacks a bond between the C2 and C3 carbons (e.g., UNA). In some embodiments, the sugar-modified nucleoside comprises a bicyclohexose nucleic acid or a tricyclic nucleic acid. In some embodiments, the modified nucleoside comprises a nucleoside in which the sugar moiety is replaced with a non-sugar moiety, e.g., a peptide nucleic acid (PNA) or a morpholino nucleic acid.

[0134] In some embodiments, the guide RNA comprises one or more modified sugars. In some embodiments, sugar modifications include modifications made by altering the substituent on the ribose ring to a group other than hydrogen or to the 2'-OH group naturally found in DNA and RNA nucleosides. In some embodiments, the substituent is introduced at the 2', 3', 4', or 5' position, or a combination thereof. In some embodiments, the nucleoside having a modified sugar moiety comprises a 2' modified nucleoside, e.g., a 2' substituted nucleoside. A 2' sugar modified nucleoside, in some embodiments, is a nucleoside having a substituent other than -H or -OH at the 2' position (2' substituted nucleoside) or comprises a 2' linked biradical, and includes 2' substituted nucleosides and LNA (2'-4' biradical bridged) nucleosides. Examples of 2'-substituted modified nucleosides include, but are not limited to, 2'-O-alkyl-RNA, 2'-O-methyl-RNA, 2'-alkoxy-RNA, 2'-O-methoxyethyl-RNA (MOE), 2'-amino-DNA, 2'-fluoro-RNA, and 2'-F-ANA nucleosides. In some embodiments, the modification in the ribose group comprises a modification at the 2' position of the ribose group. In some embodiments, the modification at the 2' position of the ribose group is selected from the group consisting of 2'-O-methyl, 2'-fluoro, 2'-deoxy, and 2'-O-(2-methoxyethyl).

[0135] In some embodiments, the guide RNA comprises one or more modified sugars. In some embodiments, the guide RNA comprises only modified sugars. In certain embodiments, the guide RNA comprises more than about 10%, 25%, 50%, 75%, or 90% modified sugars. In some embodiments, the modified sugar is a bicyclic sugar. In some embodiments, the modified sugar comprises a 2'-O-methoxyethyl group. In some embodiments, the guide RNA comprises both an internucleoside linker modification and a nucleoside modification.

[0136] In some cases, the guide RNA comprises a sequence complementary to a eukaryotic, fungal, plant, mammalian, or human genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence complementary to a eukaryotic genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence complementary to a fungal genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence complementary to a plant genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence complementary to a mammalian genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence complementary to a human genomic polynucleotide sequence.

[0137] In some embodiments, the guide RNA is 30-250 nucleotides in length. In some embodiments, the guide RNA is more than 90 nucleotides in length. In some embodiments, the guide RNA is less than 245 nucleotides in length. In some embodiments, the guide RNA is 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 220, 240, or more than 240 nucleotides in length. In some embodiments, the guide RNA is about 30 to about 40, about 30 to about 50, about 30 to about 60, about 30 to about 70, about 30 to about 80, about 30 to about 90, about 30 to about 100, about 30 to about 120, about 30 to about 140, about 30 to about 160, about 30 to about 180, about 30 to about 200, about 30 to about 220, about 30 to about 240, about 50 to about 60, about 50 to about 70, about 50 to about 80, about 50 to about 90, about 50 to about 10 ... The length is about 120, about 50 to about 140, about 50 to about 160, about 50 to about 180, about 50 to about 200, about 50 to about 220, about 50 to about 240, about 100 to about 120, about 100 to about 140, about 100 to about 160, about 100 to about 180, about 100 to about 200, about 100 to about 220, about 100 to about 240, about 160 to about 180, about 160 to about 200, about 160 to about 220, or about 160 to about 240 nucleotides.

[0138] In some cases, the left-hand recombinase sequence comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 70% identity to SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 75% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 80% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 85% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 90% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 91% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 92% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 93% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 94% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 95% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78.In some cases, the left-hand recombinase sequence comprises a sequence having at least about 96% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 97% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 98% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 99% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having 100% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78.

[0139] In some cases, the right recombinase sequence comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 70% identity to SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 75% identity to SEQ ID NO:8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 80% identity to SEQ ID NO:8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 85% identity to SEQ ID NO:8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 90% identity to SEQ ID NO:8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 91% identity to SEQ ID NO:8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 92% identity to SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 93% identity to SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 94% identity to SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93.In some cases, the right recombinase sequence comprises a sequence having at least about 95% identity to SEQ ID NO: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 96% identity to SEQ ID NO: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 97% identity to SEQ ID NO: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 98% identity to SEQ ID NO: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 99% identity to SEQ ID NO: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having 100% identity to SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93.

[0140] In some cases, the class 2 V-type Cas effector and the Tn7-type transposase complex are encoded by a polynucleotide sequence that comprises less than about 20 kilobases, less than about 15 kilobases, less than about 10 kilobases, or less than about 5 kilobases.

[0141] In some embodiments, the class 2 type V effector comprises a nuclear localization sequence (NLS). In some embodiments, the NLS is at the N-terminus of the class 2 type V effector. In some embodiments, the NLS is at the C-terminus of the class 2 type V effector. In some embodiments, the NLS is at both the N-terminus and the C-terminus of the class 2 type V effector.

[0142] In some embodiments, the NLS comprises any one of SEQ ID NOs: 192-207, or a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 192-207. In some cases, the NLS comprises a sequence having at least about 80% identity to SEQ ID NOs: 192-207. In some cases, the NLS comprises a sequence having at least about 85% identity to SEQ ID NOs: 192-207. In some cases, the NLS comprises a sequence having at least about 90% identity to SEQ ID NO: 192-207. In some cases, the NLS comprises a sequence having at least about 91% identity to SEQ ID NO: 192-207. In some cases, the NLS comprises a sequence having at least about 92% identity to SEQ ID NO: 192-207. In some cases, the NLS comprises a sequence having at least about 93% identity to SEQ ID NO: 192-207. In some cases, the NLS comprises a sequence having at least about 94% identity to SEQ ID NO: 192-207. In some cases, the NLS comprises a sequence having at least about 95% identity to SEQ ID NO: 192-207. In some cases, the NLS comprises a sequence having at least about 96% identity to SEQ ID NO: 192-207. In some cases, the NLS comprises a sequence having at least about 97% identity to SEQ ID NO: 192-207. In some cases, the NLS comprises a sequence having at least about 98% identity to SEQ ID NO: 192-207. In some cases, the NLS comprises a sequence having at least about 99% identity to SEQ ID NO: 192-207. In some cases, the NLS comprises a sequence having 100% identity to SEQ ID NO: 192-207.

[0143]

Table 1

[0144] In some embodiments, the Cas effector complex further comprises a small prokaryotic ribosomal protein subunit, S15. In some embodiments, the S15 fusion protein is encoded by a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 181-183. In some cases, S15 is encoded by a sequence having at least about 70% identity to SEQ ID NOs: 181-183. In some cases, S15 is encoded by a sequence having at least about 75% identity to SEQ ID NO: 181-183. In some cases, S15 is encoded by a sequence having at least about 80% identity to SEQ ID NO: 181-183. In some cases, S15 is encoded by a sequence having at least about 85% identity to SEQ ID NO: 181-183. In some cases, S15 is encoded by a sequence having at least about 90% identity to SEQ ID NO: 181-183. In some cases, S15 is encoded by a sequence having at least about 91% identity to SEQ ID NO: 181-183. In some cases, S15 is encoded by a sequence having at least about 92% identity to SEQ ID NO: 181-183. In some cases, S15 is encoded by a sequence having at least about 93% identity to SEQ ID NO: 181-183. In some cases, S15 is encoded by a sequence having at least about 94% identity to SEQ ID NOs: 181-183. In some cases, S15 is encoded by a sequence having at least about 95% identity to SEQ ID NOs: 181-183. In some cases, S15 is encoded by a sequence having at least about 96% identity to SEQ ID NOs: 181-183.In some cases, S15 is encoded by a sequence having at least about 97% identity to SEQ ID NOs: 181-183. In some cases, S15 is encoded by a sequence having at least about 98% identity to SEQ ID NOs: 181-183. In some cases, S15 is encoded by a sequence having at least about 99% identity to SEQ ID NOs: 181-183. In some cases, S15 is encoded by a sequence having 100% identity to SEQ ID NOs: 181-183.

[0145] In some cases, S15 comprises a sequence having at least about 70% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 75% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 80% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 85% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 90% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 91% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 92% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 93% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 94% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 95% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 96% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 97% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 98% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 99% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having 100% identity to SEQ ID NO: 187-189.

[0146] In some embodiments, the Cas effector complex comprises one or more linkers linking the class 2 type V effector, the small prokaryotic ribosomal protein subunit S15, the transposase, the gRNA, or combinations thereof. In some embodiments, the linker comprises at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, or 400 amino acids. In some embodiments, the linker comprises at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides. In some embodiments, the linker is encoded by the sequence of SEQ ID NO: 186 or a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 186. In some embodiments, the linker is encoded by SEQ ID NO:186.

[0147] Fusion proteins In some embodiments, described herein is a system for translocating a cargo nucleotide sequence into a target nucleic acid site comprising a fusion protein or a nucleic acid encoding the fusion protein. In some embodiments, the fusion protein or the nucleic acid encoding the fusion protein comprises a class 2 type V effector, a small prokaryotic ribosomal protein subunit S15, a transposase, a gRNA, or a combination thereof. In some embodiments, the fusion protein comprises one or more transposases.

[0148] In some embodiments, a nuclear localization sequence (NLS) is fused to the class 2 type V effector. In some embodiments, the NLS is fused at the N-terminus of the class 2 type V effector. In some embodiments, the NLS is fused at the C-terminus of the class 2 type V effector. In some embodiments, the NLS is fused at both the N-terminus and the C-terminus of the class 2 type V effector.

[0149] In some embodiments, the NLS comprises any one of SEQ ID NOs: 192-207, or a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 192-207. In some cases, the NLS comprises a sequence having at least about 80% identity to SEQ ID NOs: 192-207. In some cases, the NLS comprises a sequence having at least about 85% identity to SEQ ID NOs: 192-207. In some cases, the NLS comprises a sequence having at least about 90% identity to SEQ ID NO: 192-207. In some cases, the NLS comprises a sequence having at least about 91% identity to SEQ ID NO: 192-207. In some cases, the NLS comprises a sequence having at least about 92% identity to SEQ ID NO: 192-207. In some cases, the NLS comprises a sequence having at least about 93% identity to SEQ ID NO: 192-207. In some cases, the NLS comprises a sequence having at least about 94% identity to SEQ ID NO: 192-207. In some cases, the NLS comprises a sequence having at least about 95% identity to SEQ ID NO: 192-207. In some cases, the NLS comprises a sequence having at least about 96% identity to SEQ ID NO: 192-207. In some cases, the NLS comprises a sequence having at least about 97% identity to SEQ ID NO: 192-207. In some cases, the NLS comprises a sequence having at least about 98% identity to SEQ ID NO: 192-207. In some cases, the NLS comprises a sequence having at least about 99% identity to SEQ ID NO: 192-207. In some cases, the NLS comprises a sequence having 100% identity to SEQ ID NO: 192-207.

[0150] In some embodiments, the fusion protein or a nucleic acid encoding the fusion protein comprises a fusion of S15 and a nuclear localization sequence (NLS). In some embodiments, the NLS is fused at the N-terminus of S15. In some embodiments, the NLS is fused at the C-terminus of S15. In some embodiments, the NLS is fused at both the N-terminus and the C-terminus of S15.

[0151] In some embodiments, the S15 fusion protein further comprises a cleavable peptide, hi some embodiments, the peptide is a 2A peptide.

[0152] In some embodiments, the S15 fusion protein is encoded by a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 181-183. In some embodiments, the S15 fusion protein is encoded by a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 181-183. In some embodiments, the Cas effector complex further comprises a small prokaryotic ribosomal protein subunit S15. In some embodiments, the S15 fusion protein is encoded by a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 181-183. In some embodiments, the S15 fusion protein is encoded by a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 181-183. In some cases, S15 is encoded by a sequence having at least about 70% identity to SEQ ID NOs: 181-183. In some cases, S15 is encoded by a sequence having at least about 75% identity to SEQ ID NOs: 181-183. In some cases, S15 is encoded by a sequence having at least about 80% identity to SEQ ID NOs: 181-183.In some cases, S15 is encoded by a sequence having at least about 85% identity to SEQ ID NO: 181-183. In some cases, S15 is encoded by a sequence having at least about 90% identity to SEQ ID NO: 181-183. In some cases, S15 is encoded by a sequence having at least about 91% identity to SEQ ID NO: 181-183. In some cases, S15 is encoded by a sequence having at least about 92% identity to SEQ ID NO: 181-183. In some cases, S15 is encoded by a sequence having at least about 93% identity to SEQ ID NO: 181-183. In some cases, S15 is encoded by a sequence having at least about 94% identity to SEQ ID NO: 181-183. In some cases, S15 is encoded by a sequence having at least about 95% identity to SEQ ID NO: 181-183. In some cases, S15 is encoded by a sequence having at least about 96% identity to SEQ ID NO: 181-183. In some cases, S15 is encoded by a sequence having at least about 97% identity to SEQ ID NO: 181-183. In some cases, S15 is encoded by a sequence having at least about 98% identity to SEQ ID NO: 181-183. In some cases, S15 is encoded by a sequence having at least about 99% identity to SEQ ID NO: 181-183. In some cases, S15 is encoded by a sequence having 100% identity to SEQ ID NO: 181-183.

[0153] In some embodiments, the S15 fusion protein comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 187-189. In some embodiments, the S15 fusion protein comprises at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 187-189. In some cases, S15 comprises a sequence having at least about 70% identity to SEQ ID NOs: 187-189. In some cases, S15 comprises a sequence having at least about 75% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 80% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 85% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 90% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 91% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 92% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 93% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 94% identity to SEQ ID NO: 187-189. In some cases, S15 comprises a sequence having at least about 95% identity to SEQ ID NOs: 187-189. In some cases, S15 comprises a sequence having at least about 96% identity to SEQ ID NOs: 187-189. In some cases, S15 comprises a sequence having at least about 97% identity to SEQ ID NOs: 187-189.In some cases, S15 comprises a sequence having at least about 98% identity to SEQ ID NOs: 187-189. In some cases, S15 comprises a sequence having at least about 99% identity to SEQ ID NOs: 187-189. In some cases, S15 comprises a sequence having 100% identity to SEQ ID NOs: 187-189.

[0154] In some embodiments, the NLS is fused to a transposase. In some embodiments, the transposase is TnsB, TnsC, or TniQ. In some embodiments, the transposase is TnsB. In some embodiments, the transposase is TnsC. In some embodiments, the transposase is TniQ. In some embodiments, the NLS is fused at the N-terminus of the transposase. In some embodiments, the NLS is fused at the C-terminus of the transposase. In some embodiments, the NLS is fused at both the N-terminus and the C-terminus of the transposase.

[0155] In some embodiments, the fusion protein or a nucleic acid encoding the fusion protein comprises a gRNA described herein (e.g., a dual gRNA or a single gRNA).

[0156] In some embodiments, the class 2 type V effector, small prokaryotic ribosomal protein subunit S15, transposase, gRNA, or fusion protein comprises a tag. In some embodiments, the tag is an affinity tag. In some embodiments, the tag is a polypeptide or a polynucleotide. Exemplary affinity tags include, but are not limited to, His tag, Flag tag, Myc tag, MBP tag, and GST tag.

[0157] In some embodiments, the class 2 type V effector, small prokaryotic ribosomal protein subunit S15, transposase, or fusion protein comprises a protease cleavage site. Exemplary protease cleavage sites include, but are not limited to, a TEV site, a C3 site, a factor Xa site, and an enterokinase site.

[0158] cell In certain embodiments, cells comprising the systems described herein are described herein.

[0159] In some embodiments, the cell is a eukaryotic cell (e.g., a plant cell, an animal cell, a protist cell, or a fungal cell), a mammalian cell (Chinese hamster ovary (CHO) cell, baby hamster kidney (BHK), human embryonic kidney (HEK), mouse myeloma (NS0), or a human retinal cell), an immortalized cell (e.g., a HeLa cell, a COS cell, a HEK-293T cell, an MDCK cell, a 3T3 cell, a PC12 cell, a Huh7 cell, a HepG2 cell, a K562 cell, a N2a cell, or a SY5Y cell), an insect cell (e.g., a Spodoptera frugiperda cell, a Trichoplusia ni cell, a Drosophila melanogaster cell, a S2 cell, or a Heliothis virescens cell), a yeast cell (e.g., a Saccharomyces In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is an immortalized cell. In some embodiments, the cell is an insect cell. In some embodiments, the cell is a yeast cell. In some embodiments, the cell is a plant cell. In some embodiments, the cell is a fungal cell. In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is a yeast cell. In some embodiments, the cell is a plant cell. In some embodiments, the cell is a fungal cell. In some embodiments, the cell is a prokaryotic cell.

[0160] In some embodiments, the cells are A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cells, HT1080, HepG2, Huh7, K562, primary cells, or derivatives thereof.

[0161] Delivery and Vectors In some embodiments, disclosed herein are nucleic acid sequences encoding the MG64 system, including a class 2 type V effector, a small prokaryotic ribosomal protein subunit S15, a transposase, a gRNA, a fusion protein, or a gene editing system disclosed herein.

[0162] In some embodiments, the nucleic acid encoding the MG64 system is DNA, e.g., linear DNA, plasmid DNA, or minicircle DNA. In some embodiments, the nucleic acid encoding the MG64 system is RNA, e.g., mRNA.

[0163] In some embodiments, the nucleic acid encoding the MG64 system is delivered by a nucleic acid-based vector. In some embodiments, the nucleic acid-based vector is a plasmid (e.g., a circular DNA molecule that can replicate autonomously inside a cell), a cosmid (e.g., a pWE or sCos vector), an artificial chromosome, a human artificial chromosome (HAC), a yeast artificial chromosome (YAC), a bacterial artificial chromosome (BAC), a P1-derived artificial chromosome (PAC), a phagemid, a phage derivative, a bacmid, or a virus. In some embodiments, the nucleic acid based vector is pSF-CMV-NEO-NH2-PPT-3XFLAG, pSF-CMV-NEO-COOH-3XFLAG, pSF-CMV-PURO-NH2-GST-TEV, pSF-OXB20-COOH-TEV-FLAG(R)-6His, pCEP4 pDEST27, pSF-CMV-Ub-KrYFP, pSF-CMV-FMDV-daGFP, pEF1a-mCherry-N1 vector, pEF1a-tdTomato vector, pSF-CMV-FMDV-Hygro, pSF-CMV-PGK-Puro, pMCP-tag(m), pSF-CMV-PURO-NH2-CMYC, pSF-OXB20-BetaGal, pSF-OXB20-Fluc, pSF-OXB20, pSF-Tac, pRI 101-AN The vector is selected from the list consisting of pCambia2301, pTYB21, pKLAC2, pAc5.1 / V5-His A, and pDEST8.

[0164] In some embodiments, the nucleic acid-based vector comprises a promoter. In some embodiments, the promoter is selected from the group consisting of a minipromoter, an inducible promoter, a constitutive promoter, and derivatives thereof. In some embodiments, the promoter is selected from the group consisting of CMV, CBA, EF1a, CAG, PGK, TRE, U6, UAS, T7, Sp6, lac, araBad, trp, Ptac, p5, p19, p40, synapsin, CaMKII, GRK1, and derivatives thereof. In some embodiments, the promoter is a U6 promoter. In some embodiments, the promoter is a CAG promoter. In some embodiments, the promoter is encoded by any one of SEQ ID NOs: 190-191, or a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 190-191.

[0165] In some embodiments, the nucleic acid based vector is a virus. In some embodiments, the virus is an alphavirus, parvovirus, adenovirus, AAV, baculovirus, dengue virus, lentivirus, herpes virus, poxvirus, anellovirus, bocavirus, vaccinia virus, or retrovirus. In some embodiments, the virus is an alphavirus. In some embodiments, the virus is a parvovirus. In some embodiments, the virus is an adenovirus. In some embodiments, the virus is an AAV. In some embodiments, the virus is a baculovirus. In some embodiments, the virus is a dengue virus. In some embodiments, the virus is a lentivirus. In some embodiments, the virus is a herpes virus. In some embodiments, the virus is a poxvirus. In some embodiments, the virus is anellovirus. In some embodiments, the virus is a bocavirus. In some embodiments, the virus is a vaccinia virus. In some embodiments, the virus is a retrovirus.

[0166] In some embodiments, the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV-rh8, AAV-rh 10, AAV-rh20, AAV-rh39, AAV-rh74, AAV-rhM4-1, AAV-hu37, AAV-Anc80, AAV-Anc80L65, AAV-7m8, AAV-PHP-B, AAV-PHP-EB, AAV-2.5, AAV-2tYF, AAV-3B, AAV-LK03, AAV-HSC1, AAV-HSC2, AAV-HSC3, AAV-HSC4, AAV-HSC5, AAV-HSC6, AAV-HSC7, AAV-HSC8, AAV-HSC9, AAV-HSC10, AAV-HSC11, AAV-HSC12, AAV-HSC13, AAV-HSC14, AAV-HSC15, AAV-TT, AAV-DJ / 8, AAV-Myo, AAV-NP40, AAV-NP59, AAV-NP22, AAV-NP66, AAV-HSC16, or derivatives thereof. In some embodiments, the herpes virus is HSV type 1, HSV-2, VZV, EBV, CMV, HHV-6, HHV-7, or HHV-8.

[0167] In some embodiments, the virus is AAV1 or a derivative thereof. In some embodiments, the virus is AAV2 or a derivative thereof. In some embodiments, the virus is AAV3 or a derivative thereof. In some embodiments, the virus is AAV4 or a derivative thereof. In some embodiments, the virus is AAV5 or a derivative thereof. In some embodiments, the virus is AAV6 or a derivative thereof. In some embodiments, the virus is AAV7 or a derivative thereof. In some embodiments, the virus is AAV8 or a derivative thereof. In some embodiments, the virus is AAV9 or a derivative thereof. In some embodiments, the virus is AAV10 or a derivative thereof. In some embodiments, the virus is AAV11 or a derivative thereof. In some embodiments, the virus is AAV12 or a derivative thereof. In some embodiments, the virus is AAV13 or a derivative thereof. In some embodiments, the virus is AAV14 or a derivative thereof. In some embodiments, the virus is AAV15 or a derivative thereof. In some embodiments, the virus is AAV16 or a derivative thereof. In some embodiments, the virus is AAV-rh8 or a derivative thereof. In some embodiments, the virus is AAV-rh10 or a derivative thereof. In some embodiments, the virus is AAV-rh20 or a derivative thereof. In some embodiments, the virus is AAV-rh39 or a derivative thereof. In some embodiments, the virus is AAV-rh74 or a derivative thereof. In some embodiments, the virus is AAV-rhM4-1 or a derivative thereof. In some embodiments, the virus is AAV-hu37 or a derivative thereof. In some embodiments, the virus is AAV-Anc80 or a derivative thereof. In some embodiments, the virus is AAV-Anc80L65 or a derivative thereof. In some embodiments, the virus is AAV-7m8 or a derivative thereof. In some embodiments, the virus is AAV-PHP-B or a derivative thereof. In some embodiments, the virus is AAV-PHP-EB or a derivative thereof.In some embodiments, the virus is AAV-2.5 or a derivative thereof. In some embodiments, the virus is AAV-2tYF or a derivative thereof. In some embodiments, the virus is AAV-3B or a derivative thereof. In some embodiments, the virus is AAV-LK03 or a derivative thereof. In some embodiments, the virus is AAV-HSC1 or a derivative thereof. In some embodiments, the virus is AAV-HSC2 or a derivative thereof. In some embodiments, the virus is AAV-HSC3 or a derivative thereof. In some embodiments, the virus is AAV-HSC4 or a derivative thereof. In some embodiments, the virus is AAV-HSC5 or a derivative thereof. In some embodiments, the virus is AAV-HSC6 or a derivative thereof. In some embodiments, the virus is AAV-HSC7 or a derivative thereof. In some embodiments, the virus is AAV-HSC8 or a derivative thereof. In some embodiments, the virus is AAV-HSC9 or a derivative thereof. In some embodiments, the virus is AAV-HSC10 or a derivative thereof. In some embodiments, the virus is AAV-HSC11 or a derivative thereof. In some embodiments, the virus is AAV-HSC12 or a derivative thereof. In some embodiments, the virus is AAV-HSC13 or a derivative thereof. In some embodiments, the virus is AAV-HSC14 or a derivative thereof. In some embodiments, the virus is AAV-HSC15 or a derivative thereof. In some embodiments, the virus is AAV-TT or a derivative thereof. In some embodiments, the virus is AAV-DJ / 8 or a derivative thereof. In some embodiments, the virus is AAV-Myo or a derivative thereof. In some embodiments, the virus is AAV-NP40 or a derivative thereof. In some embodiments, the virus is AAV-NP59 or a derivative thereof. In some embodiments, the virus is AAV-NP22 or a derivative thereof. In some embodiments, the virus is AAV-NP66 or a derivative thereof. In some embodiments, the virus is AAV-HSC16 or a derivative thereof.

[0168] In some embodiments, the virus is HSV-1 or a derivative thereof. In some embodiments, the virus is HSV-2 or a derivative thereof. In some embodiments, the virus is VZV or a derivative thereof. In some embodiments, the virus is EBV or a derivative thereof. In some embodiments, the virus is CMV or a derivative thereof. In some embodiments, the virus is HHV-6 or a derivative thereof. In some embodiments, the virus is HHV-7 or a derivative thereof. In some embodiments, the virus is HHV-8 or a derivative thereof.

[0169] In some embodiments, the nucleic acid encoding the MG64 system is delivered by a non-nucleic acid based delivery system (e.g., a non-viral delivery system). In some embodiments, the non-viral delivery system is a liposome. In some embodiments, the nucleic acid is associated with a lipid. The nucleic acid associated with a lipid is in some embodiments encapsulated in the aqueous interior of the liposome, interspersed within the lipid bilayer of the liposome, attached to the liposome via a linking molecule associated with both the liposome and the nucleic acid, entrapped in the liposome, complexed with the liposome, dispersed in a solution containing lipid, mixed with lipid, combined with lipid, contained as a suspension in lipid, contained in or complexed with micelles, or otherwise associated with lipid. In some embodiments, the nucleic acid is included in a lipid nanoparticle (LNP).

[0170] In some embodiments, the fusion protein or genome editing system is introduced into the cell in any suitable manner, either stably or transiently. In some embodiments, the fusion protein or genome editing system is transfected into the cell. In some embodiments, the cell is transduced or transfected with a nucleic acid construct encoding the fusion protein or genome editing system. For example, the cell is transduced (e.g., with a virus encoding the fusion protein or genome editing system) or transfected with a nucleic acid encoding the fusion protein or genome editing system, or a translated fusion protein or genome editing system (e.g., with a plasmid encoding the fusion protein or genome editing system). In some embodiments, the transduction is stable or transient transduction. In some embodiments, the cell expressing or containing the fusion protein or genome editing system is transduced or transfected with one or more gRNA molecules, for example, when the fusion protein or genome editing system contains a CRISPR nuclease. In some embodiments, a plasmid expressing a fusion protein or a genome editing system is introduced into a cell through electroporation, transient (e.g., lipofection), stable genome integration (e.g., piggybac), and viral transduction (e.g., lentivirus or AAV), or other methods known to those skilled in the art. In some embodiments, the gene editing system is introduced into a cell as one or more polypeptides. In some embodiments, delivery is achieved through the use of an RNP complex. Methods for delivering polypeptides and / or RNPs into cells are known in the art, for example, by electroporation or by cell squeezing.

[0171] Exemplary methods of nucleic acid delivery include lipofection, nucleofection, electroporation, stable genome integration (e.g., piggybac), microinjection, biolistec, virosomes, liposomes, immunoliposomes, polycations or lipid nucleic acid conjugates, naked DNA, artificial virions, and drug-enhanced uptake of DNA.Lipofection is described, for example, in U.S. Pat. Nos. 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam™, Lipofectin™, and SF Cell Line 4D-Nucleofector X Kit™ (Lonza)).Cationic and neutral lipids that are suitable for efficient receptor-recognition lipofection of polynucleotides include the lipids of WO91 / 17424 and WO91 / 16024. In some embodiments, delivery is to a cell (e.g., in vitro or ex vivo administration) or to a target tissue (e.g., in vivo administration). In some embodiments, the nucleic acid is contained in a liposome or nanoparticle that specifically targets the host cell.

[0172] Additional methods for delivery of nucleic acids into cells are known to those of skill in the art, see, e.g., US2003 / 0087817.

[0173] In some embodiments, the disclosure provides a cell comprising a vector or nucleic acid described herein. In some embodiments, the cell expresses a gene editing system or a portion thereof. In some embodiments, the cell is a human cell. In some embodiments, the cell is genome edited ex vivo. In some embodiments, the cell is genome edited in vivo.

[0174] Methods for the rearrangement The present disclosure provides a method for translocating a cargo nucleotide sequence into a target nucleic acid site. In some embodiments, the method comprises expressing a system described herein in a cell or introducing a system described herein into a cell. In some embodiments, the method comprises contacting a cell with a system described herein.

[0175] In some embodiments, the method comprises contacting a double-stranded nucleic acid comprising a cargo nucleotide sequence with a Cas effector complex comprising a class 2 V-type Cas effector and at least one engineered guide polynucleotide configured to hybridize to a target nucleotide sequence. In some embodiments, the method comprises contacting a double-stranded nucleic acid comprising the cargo nucleotide sequence with a Tn7-type transposase complex configured to bind to the Cas effector complex, the Tn7-type transposase complex comprising a TnsB subunit. In some embodiments, the method comprises contacting a double-stranded nucleic acid comprising the cargo nucleotide sequence with a double-stranded target nucleic acid comprising a target nucleic acid site.

[0176] In some cases, the cargo nucleotide sequence is adjacent to a left transposase recognition sequence. In some cases, the cargo nucleotide sequence is adjacent to a right transposase recognition sequence. In some cases, the cargo nucleotide sequence is adjacent to a left transposase recognition sequence and a right transposase recognition sequence. In some embodiments, the method further comprises a PAM sequence compatible with the nuclease adjacent to the target nucleic acid site. In some cases, the PAM sequence is located 3' to the target nucleic acid site.

[0177] In some cases, the engineered guide polynucleotide is configured to bind to a Class 2 type V Cas endonuclease. In some cases, a Class 2 type V Cas effector comprises a polypeptide comprising a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 70% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 75% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 80% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 85% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least about 90% identity to SEQ ID NO:1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least about 91% identity to SEQ ID NO:1, 12, 16, 20-30, 64, 80-85, and 220.In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 92% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 93% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 94% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 95% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 96% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 97% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 98% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, class 2 V-type Cas effectors include polypeptides comprising sequences having at least about 99% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having 100% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220.

[0178] In some cases, the TnsB subunit comprises a polypeptide having a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsA subunit comprises a polypeptide having a sequence identical to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 70% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 75% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 80% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 85% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 90% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 91% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 92% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 93% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 94% identity to SEQ ID NOs: 2, 13, 17, and 65.In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 95% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 96% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 97% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 98% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 99% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having 100% identity to SEQ ID NOs: 2, 13, 17, and 65.

[0179] In some cases, the Tn7-type transposase complex comprises at least one polypeptide (e.g., at least one, two, three, four, five, six, or more than six polypeptides) comprising a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs:3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 70% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 75% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 80% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 85% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 90% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 91% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 92% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 93% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 94% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 95% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 96% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 97% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 98% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having at least about 99% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises a polypeptide comprising a sequence having 100% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.

[0180] In some cases, the Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs:3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least 70% sequence identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, the Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently having at least about 85% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least about 90% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least about 91% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least about 92% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least about 93% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least about 94% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least about 95% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least about 96% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least about 97% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least about 98% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least about 99% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising at least 100% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.

[0181] In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 70% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 75% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 80% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 85% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 90% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222.In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 91% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 92% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 93% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 94% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 95% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 96% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 97% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 98% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222.In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 99% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having 100% identity to SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222.

[0182] In some cases, the left-hand recombinase sequence comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase comprises a sequence having at least about 70% identity to SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 75% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 80% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 85% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 90% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 91% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 92% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 93% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 94% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 95% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78.In some cases, the left-hand recombinase sequence comprises a sequence having at least about 96% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 97% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 98% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 99% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having 100% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78.

[0183] In some cases, the right recombinase sequence comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 70% identity to SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 75% identity to SEQ ID NO:8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 80% identity to SEQ ID NO:8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 85% identity to SEQ ID NO:8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 90% identity to SEQ ID NO:8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 91% identity to SEQ ID NO:8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 92% identity to SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 93% identity to SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 94% identity to SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93.In some cases, the right recombinase sequence comprises a sequence having at least about 95% identity to SEQ ID NO: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 96% identity to SEQ ID NO: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 97% identity to SEQ ID NO: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 98% identity to SEQ ID NO: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 99% identity to SEQ ID NO: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having 100% identity to SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93.

[0184] In some cases, the class 2 V-type Cas effector and the Tn7-type transposase complex are encoded by a polynucleotide sequence that comprises less than about 20 kilobases, less than about 15 kilobases, less than about 10 kilobases, or less than about 5 kilobases.

[0185] use The disclosed system can be used for a variety of applications, such as, for example, nucleic acid editing (e.g., gene editing) or binding (e.g., sequence-specific binding) to nucleic acid molecules. Such systems can be used, for example, to correct (e.g., remove or replace) genetically inherited mutations that may cause disease in a subject, to inactivate genes to confirm their function in cells, as diagnostic tools to detect disease-causing genetic elements (e.g., via cleavage of reverse-transcribed viral RNA or amplified DNA sequences encoding disease-causing mutations), as inactivated enzymes combined with probes to target and detect specific nucleotide sequences (e.g., sequences encoding antibiotic resistance in bacteria), to inactivate viruses by targeting viral genomes or to render them unable to infect host cells, to add genes or modify metabolic pathways to engineer organisms to produce valuable small molecules, macromolecules, or secondary metabolites, to establish gene drive elements for evolutionary selection, and / or to detect cellular perturbations by exogenous small molecules and nucleotides as biosensors.

[0186] kit In some embodiments, the disclosure provides kits that include one or more nucleic acid constructs encoding various components of the fusion proteins or genome editing systems described herein, including, for example, nucleotide sequences encoding components of a fusion protein or genome editing system capable of modifying a target DNA sequence. In some embodiments, the nucleotide sequences include a heterologous promoter that drives expression of an RNA genome editing system component.

[0187] In some embodiments, the fusion protein or gene editing system comprising the class 2 type V effector, small prokaryotic ribosomal protein subunit S15, transposase, gRNA, or any combination thereof disclosed herein is incorporated into a pharmaceutical, diagnostic, or research kit to facilitate its use in therapeutic, diagnostic, or research applications. The kit may include one or more containers housing any of the vectors disclosed herein, and instructions for use.

[0188] The kits can be designed to facilitate the use of the methods described herein by researchers and can take many forms. Each of the compositions of the kit can be provided in liquid form (e.g., in solution) or in solid form (e.g., dry powder), if applicable. In certain cases, some of the compositions may be configurable or otherwise processable (e.g., into an active form), for example, by the addition of suitable solvents or other species (e.g., water or cell culture medium), which may or may not be provided with the kit. As used herein, "instructions" defines an instructional and / or promotional component, and can typically involve written instructions on or associated with the packaging of the present disclosure. Instructions can also include any verbal or electronic instructions provided in any manner, such as audiovisual (e.g., videotape, DVD, etc.), internet, and / or web-based communication, etc., such that the user clearly recognizes that the instructions are related to the kit. The written instructions, in some embodiments, are in a form prescribed by a government agency that regulates the manufacture, use, or sale of pharmaceutical or biological products, and the instructions may also reflect approval by the agency of the manufacture, use, or sale for animal administration. EXAMPLES

[0189] The following examples are given for the purpose of illustrating various embodiments of the present disclosure, and are not intended to limit the present disclosure in any manner. The examples, together with the methods described herein, are currently representative of preferred embodiments, are exemplary, and are not intended as limitations on the scope of the present disclosure. Modifications therein and other uses encompassed within the spirit of the present disclosure as defined by the scope of the claims will occur to those skilled in the art.

[0190] Example 1 - (General Protocol) PAM Sequence Identification / Confirmation of the System Described Herein Putative endonucleases were expressed in an E. coli lysate-based expression system. The PAM sequences were determined by sequencing a plasmid containing randomly generated potential PAM sequences that could be cleaved by the putative nuclease. In this system, E. coli codon-optimized nucleotide sequences encoding the putative nucleases were transcribed and translated in vitro from a PCR fragment under the control of a T7 promoter. A second PCR fragment carrying a T7 promoter followed by a minimal CRISPR array composed of a repeat-spacer-repeat sequence was transcribed in the same reaction. Successful expression of the endonuclease sequence and the repeat-spacer-repeat sequence in an E. coli lysate-based expression system, followed by CRISPR array processing, provided an active in vitro CRISPR nuclease complex.

[0191] A library of target plasmids containing spacer sequences matching those in the minimal array preceded by 8N mixed bases (potential PAM sequences) was incubated with the products of the expression system reaction. After 1–3 h, the reaction was stopped and DNA was recovered via a DNA clean-up kit. Adapter sequences were blunt-end ligated to DNA with active PAM sequences cleaved by the endonuclease, while uncleaved DNA was inaccessible for ligation. DNA segments containing active PAM sequences were then amplified by PCR using primers specific for the library and adapter sequences. PCR amplification products were resolved on a gel to identify amplicons corresponding to cleavage events. Amplified segments of the cleavage reaction were also used as templates for preparation of NGS libraries or as substrates for Sanger sequencing. Sequencing this resulting library, a subset of the starting 8N library, revealed sequences with PAM activity compatible with the CRISPR complex. For PAM testing with treated RNA constructs, the same procedure was repeated except that in vitro transcribed RNA was added along with the plasmid library and the minimal CRISPR array template was omitted.

[0192] Analysis of the intergenic regions surrounding the Cas effectors and CRISPR arrays identified potential anti-repeat sequences that correspond to the double-stranded sequence of tracrRNA. TracrRNA and the crRNA repeats were folded and trimmed, and a GAAA tetraloop sequence was added to maintain the stem-loop region of the crRNA-tracrRNA complex.

[0193] Example 2A - In vitro targeted integrase activity Integrase activity was assayed using previously identified PAMs, but could alternatively be performed with PAM library substrates of reduced efficiency. One arrangement of components for in vitro testing involved three plasmids other than the one containing the donor sequence: (1) an expression plasmid with an effector (or effectors) under a T7 promoter, (2) an expression plasmid with a transposase gene under a T7 promoter, sgRNA or crRNA and tracrRNA, (3) a target plasmid that contained a spacer site and an appropriate PAM, and (4) a donor plasmid containing the necessary left end (LE) and right end (RE) DNA sequences for transposition around a cargo gene (e.g., a selection marker such as a Tet resistance gene). An in vitro transcription / translation system (e.g., an E. coli lysate or reticulocyte lysate-based system) was used to express the effector and transposase genes. After expression, RNA, target DNA, and donor DNA were added and incubated to allow transposition to occur. Transcription was detected via PCR across the transposase site junction, with one primer on the target DNA and one on the donor DNA. The resulting PCR products were sequenced via NGS to determine the exact insertion topology relative to the sgRNA / crRNA target site. Primers were positioned downstream such that various insertion sites could be accommodated and detected. Primers were designed such that integration could be detected in either orientation of the cargo and on either side of the spacer, as integration direction has also not been previously documented.

[0194] Integration efficiency was measured via qPCR of the experimental output of target DNA with integrated cargo, normalized to the amount of unmodified target DNA, also measured via quantitative PCR (qPCR) measurement.

[0195] This assay can be performed with purified protein components rather than from lysate-based expression. In this case, proteins were expressed in E. coli protease-deficient B strain under a T7 inducible promoter, cells were lysed using sonication, and His-tagged proteins of interest were purified using Ni-NTA affinity chromatography on an FPLC system. Purity was determined using SDS-PAGE and densitometry of resolved protein bands on Coomassie-stained acrylamide gels. Proteins were desalted in a storage buffer consisting of 50 mM Tris-HCl, 300 mM NaCl, 1 mM TCEP, 5% glycerol at pH 7.5 (or other buffer determined for maximum stability) and stored at -80°C. After purification, the effector and transposase were added to the sgRNA, target DNA, and donor DNA described above in reaction buffer, e.g., 26 mM HEPES at pH 7.5, 4.2 mM TRIS at pH 8, 50 μg / mL BSA, 2 mM ATP, 2.1 mM DTT, 0.05 mM EDTA, 0.2 mM MgCl2, 28 mM NaCl, 21 mM KCl, 1.35% glycerol (final pH 7.5), supplemented with 15 mM Mg(Oac)2.

[0196] Example 2B - In vitro activity Targeted Nucleases In situ expression and protein sequence analysis showed that several RNA-guided effectors were active nucleases: they contained predicted endonuclease-associated domains (matching the RuvC and HNH endonuclease domains) and / or predicted HNH and RuvC catalytic residues.

[0197] Candidate activity was tested with engineered single guide RNA sequences using an E. coli lysate-based expression system and in vitro transcribed RNA. Active proteins that successfully cleaved the library gave rise to a band of approximately 170 bp in the gel.

[0198] DNA integration and rearrangement A transposon is predicted to be active when the genomic sequence encoding the transposon contains one or more protein sequences with transposase and / or integrase functions within the left and right ends of the transposon. Tn7 transposons as defined herein may contain the catalytic transposase TnsB, but may also contain TnsA, TnsC, TnsD, TnsE, TniQ, and / or other transposases or integrases. The transposon ends contain predicted transposase binding sites that contain direct and / or inverted repeats of 15 bp to 150 bp in length flanking the transposase protein and other "cargo" genes. Protein sequence analysis shows that the transposases contain an integrase domain, a transposase domain, and / or transposase catalytic residues, suggesting that they are active (e.g., FIG. 4A).

[0199] Targeted DNA integration Putative CRISPR-associated transposons (CASTs) contain DNA and / or RNA-targeting CRISPR nucleases or effectors, as well as proteins with predicted transposase function in the vicinity of the CRISPR array. In some systems, the nuclease is predicted to be active based on the presence of an endonuclease-associated catalytic domain and / or catalytic residues.

[0200] In some systems, effectors are predicted to be inactive based on the absence of endonuclease domains and / or catalytic residues, although they have homology with documented CRISPR effector proteins.Transposases are predicted to associate with effectors when CRISPR loci (inactive CRISPR nucleases and arrays) and transposase proteins are located within the left and right ends of predicted transposons (Figure 4A).In this case, effectors are predicted to guide DNA integration to specific genomic locations based on guide RNAs.

[0201] CAST activity was tested using five types of components: (1) Cas effector proteins expressed by an in vitro expression system, (2) a target DNA fragment or plasmid containing a target sequence and a PAM corresponding to the Cas enzyme, (3) a donor DNA fragment containing a marker or fragment of DNA flanking the LE and RE of the transposase system in the DNA fragment or plasmid, (4) any combination of transposase proteins expressed using an in vitro expression system, and (5) an engineered in vitro transcribed single guide RNA sequence. Successful transposition of the donor fragment was assayed for active systems by PCR amplification of the donor-target junction.

[0202] After performing the transposition reaction, PCR amplification of the junctions indicated that proper donor-target formation had occurred and that the transposition reaction was sg-dependent (Figure 6). PCR amplification of reactions #3 and #4 indicated that both orientations of the donor to the target had been achieved, i.e., LE closer to the PAM and RE closer to the PAM. Although both transposition orientations were achieved, there was a preference for donor integration in the target with the LE closer to the PAM, as represented by the strong bands present for reactions #4 and #5.

[0203] Sanger sequencing of the preferred orientation products was performed. Of the integrations that occurred with the LE closer to the PAM, there was a clear degradation of the sequencing chromatogram signal from either the forward or reverse orientation across the target / donor junction. This indicated that of the products oriented with the LE closer to the PAM, integration occurred over a range of nucleotides, with the major product of the LE closer to the PAM as a 61-bp integration from the PAM (Figure 7A). Sequencing of the donor product across the donor-target junction defined the composition of the essential outer border of the LE and RE sequences (Figures 7A-7B). Further investigation of the LE and RE domains may determine the inner limits of the LE and RE sequences that are essential for transposition. Sequencing of the RE on the product of the LE closer to the PAM showed a 3-bp overlap downstream of the donor RE (Figure 7B). This is due, in part, to a Tn7 transposase integration event that cut and ligated the donor fragment at the staggered cut sites. The 3 bp overlap is smaller than the expected 5 bp overlap from other Tn7 transposases.

[0204] Sanger sequencing of PCR amplification products against the 8N library of target plasmids also revealed the PAM preference of the MG64-1 effector as nGTn / nGTt on the 5' end of the spacer (Figure 7C). NGS analysis of the PAM library targets confirmed the nGTn motif specificity at the 5' end.

[0205] Example 3 - Predicted RNA folding The predicted RNA fold of the active single RNA sequence was calculated at 37° using the method of Andronescu 2007. All hairpin loop secondary structures were singly removed from the structure and compiled iteratively into a smaller single guide. In the second approach, the tracrRNA of MG64-1 was aligned to the documented Vk-type tracrRNA and the region of unique insertion was mutated from the single guide and minimized by 57 bases. Figure 12A depicts the predicted structure of the MG64-1 sgRNA. Figure 12B depicts the predicted structure of the MG64-3 sgRNA. Figure 12C depicts the predicted structure of the MG64-5 sgRNA. The color of the base corresponds to the probability of base pairing for that base, with red representing a high probability and blue representing a low probability.

[0206] Example 4 - Transposon end verification via gel shift Transposon ends were tested for TnsB binding via electrophoretic mobility shift assay (EMSA). In this case, potential LEs or REs were synthesized as DNA fragments (100-500 bp) and end-labeled with FAM via PCR using FAM-labeled primers. TnsB proteins were synthesized in an in vitro transcription / translation system. After synthesis, 1 μL of TnsB protein was added to 50 nM of labeled RE or LE in a 10 μL reaction in binding buffer (20 mM HEPES at pH 7.5, 2.5 mM Tris at pH 7.5, 10 mM NaCl, 0.0625 mM EDTA, 5 mM TCEP, 0.005% BSA, 1 ug / mL poly(dI-dC), and 5% glycerol). The ligation was incubated at 30° for 40 min, then 2 uL of 6X loading buffer (60 mM KCl, 10 mM Tris at pH 7.6, 50% glycerol) was added. Binding reactions were resolved and visualized on a 5% TBE gel. A shift in the LE or RE in the presence of TnsB was due to successful binding and indicated transposase activity (Figure 24).

[0207] Example 5 - Integrase activity in E. coli Because E. coli lacks the ability to efficiently repair genomic double-stranded DNA breaks, transformation of E. coli with agents capable of inducing double-stranded breaks in the E. coli genome results in cell death. Exploiting this phenomenon, endonuclease or effector-assisted integrase activity was tested in E. coli by recombinantly expressing either endonuclease or effector-assisted integrase and guide RNA (e.g., as determined in Example 3) in a target strain with spacer / target and PAM sequences integrated into its genomic DNA.

[0208] The engineered strains were then transformed with plasmids containing nucleases or effectors with single guide RNAs, plasmids expressing integrases and accessory genes, and plasmids containing a temperature-sensitive origin of replication with a selectable marker flanked by left-extremity (LE) and right-extremity (RE) transposon motifs for integration. Transformants induced for expression of these genes were then screened for transfer of the marker to the genomic target by selection at the restrictive temperature for plasmid replication, and marker integration within the genome was confirmed by PCR.

[0209] An unbiased approach was used to screen for off-target integration. Briefly, purified gDNA was fragmented with Tn5 transposase or sheared, and then the DNA of interest was PCR amplified using primers specific for the ligated adapter and selectable marker. The amplicons were then prepared for NGS sequencing. Analysis of the resulting sequences was trimmed from the transposon sequence and flanking sequences were mapped to the genome to determine the insertion location and to determine the off-target insertion rate.

[0210] Example 6 - Colony PCR screening for transposase activity For testing of nuclease or effector-assisted integrase activity in bacterial cells, strain MGB0032 was constructed from BL21(DE3) E. coli cells engineered to contain a target and corresponding PAM sequence specific for MG64_1. MGB0032 E. coli cells were then transformed with pJL56 (a plasmid expressing the MG64_1 effector and helper suite, ampicillin resistant) and pTCM 64_1 sg, a chloramphenicol resistant plasmid expressing a single guide RNA sequence of the engineered target of interest driven by a T7 promoter.

[0211] MGB0032 cultures containing both plasmids were then grown to saturation, diluted at least 1:10 into growth culture containing the appropriate antibiotic, and incubated at 37°C to an OD of approximately 1. Cells from this growth step were made electrocompetent and transformed with streamlined64_1 pDonor, a plasmid carrying a tetracycline resistance marker flanked by left-end (LE) and right-end (RE) transposon motifs for integration. Electroporated cells were then plated on LB-agar-ampicillin-chloramphenicol-tetracycline and allowed to recover for 2 hours on LB medium in the presence or absence of IPTG at a final concentration of 100 μM before being incubated at 37°C for 4 days. A sterile toothpick was used to sample each resulting CFU, which was mixed into water. To this solution was added Q5 High Fidelity PCR mastermix (New England Biolabs) and primers LA155 (5'-GCTCTTCCGATCTNNNNNGATGAGCGCATTGTTAGATTTCAT-3') and oJL50 (5'-AAACCGACATCGCAGGCTTC-3'). These primers flank the predicted insertion junction. The predicted product size was 609 bp. DNA amplified PCR products were visualized on a 2% agarose gel. Sanger sequencing of the PCR products confirmed the transposition event.

[0212] Example 7 - Intracellular expression / in vitro assay To test the functionality of the NLS constructs in a physiologically relevant environment, constructs cloned with active NLS-tagged CAST components were integrated into K562 cells using lentiviral transduction. Briefly, constructs cloned into lentiviral transfer plasmids were transfected into 293T cells with envelope and packaging plasmids, and virus containing supernatants were harvested from the medium after 72 hours of incubation. The virus-containing medium was then incubated with K562 cell lines containing 8 μg / mL polybrene for 72 hours, and transfected cells were then selected for co-integration using 1 μg / mL puromycin for 4 days. The selected cell lines were harvested at the end of the 4 days and differentially lysed for nuclear and cytoplasmic fractions. Subsequent fractions were then tested for transposition competence using a complementary set of in vitro expression components.

[0213] Ten million cells were centrifuged and washed once with 1x PBS, pH 7.4. The supernatant wash was aspirated completely into the cell pellet and flash frozen at -80°C for 16 hours. After thawing on ice, the cell pellet size was measured by mass and the proteins in the cell fraction were naturally extracted using the appropriate extraction volume of cell fraction and nuclear extraction reagent (NE-PER). Briefly, the cytoplasmic extraction reagent was used at 1:10 mass of cells to volume of extraction reagent. The cell suspension was mixed by vortexing and lysed with non-ionic detergent. The cells were then centrifuged at 16,000xg for 5 minutes at 4°C. The cytoplasmic extraction supernatant was then decanted and saved for in vitro testing. The nuclear extraction reagent was then added at 1:2 original cell mass to nuclear extraction reagent and incubated on ice for 1 hour with intermittent vortexing. The nuclear suspension was then centrifuged at 16,000xg for 10 min at 4°C, and the supernatant nuclear extract was decanted and tested for in vitro transposition activity. Using 4 μL of each cell and nuclear extract for each condition, in vitro transposition reactions were performed with complementary sets of in vitro expressed proteins, donor DNA, pTarget, and buffers. Evidence of transposition activity was assayed by PCR amplification of donor-target junctions.

[0214] Example 8 - Activity in mammalian cells (prospective) To demonstrate targeting and cleavage activity in mammalian cells, nuclear localization sequences are fused to the C-terminus of each of the nuclease or effector protein and integrase protein, and the fusion protein is purified. A single guide RNA targeting the genomic locus of interest is synthesized and incubated with the nuclease / effector protein to form a ribonucleoprotein complex. Cells are transfected with a plasmid containing a selectable neomycin resistance marker (NeoR) or a fluorescent marker flanked by left-end (LE) and right-end (RE) motifs, allowed to recover for 4-6 hours, and then electroporated with the nuclease RNP and integrase protein. Integration of the plasmid into the genome is quantified by counting G418-resistant colonies or by fluorescence-activated cell cytometry. Genomic DNA is extracted 72 hours after electroporation and used for preparation of NGS libraries. Off-target frequency is assayed by shearing the genome and preparing amplicons of the transposon marker and flanking DNA for NGS library preparation. At least 40 different target sites are selected to test the activity of each targeting system.

[0215] Example 9 - Targeted Nuclease Activity In situ expression and protein sequence analysis suggested that several RNA-guided effectors were active nucleases: they contained predicted endonuclease-associated domains (matching RuvC and HNH endonuclease domains) and predicted HNH and RuvC catalytic residues (Figure 4A).

[0216] Candidate activity was tested with the engineered single guide RNA sequences using an in vitro expression system and in vitro transcribed RNA. Active proteins that successfully cleaved the library gave rise to a band of approximately 170 bp in the gel.

[0217] Example 10 - Identification of transposons A transposon is predicted to be active when it contains one or more protein sequences with transposase and / or integrase functions between the left and right ends of the transposon. Tn7 transposons as defined herein may contain the catalytic transposase TnsB, but may also contain TnsA, TnsC, TnsD, TnsE, TniQ, and / or other transposases or integrases. The transposon ends contain predicted transposase binding sites that contain direct and / or inverted repeats of 15 bp to 150 bp in length that flank the transposase protein and other "cargo" genes. Protein sequence analysis shows that the transposases contain an integrase domain, a transposase domain, and / or transposase catalytic residues, suggesting that they are active (e.g., Figure 4A and panel A of Figure 5).

[0218] Example 11 - Identification of CRISPR-associated transposons Putative CRISPR-associated transposons (CASTs) contain DNA and / or RNA-targeting CRISPR effectors and proteins with predicted transposase function in the vicinity of CRISPR arrays. In some systems, effectors are predicted to have nuclease activity based on the presence of endonuclease-associated catalytic domains and / or catalytic residues (e.g., FIG. 4A). Transposases were predicted to be associated with effectors when the CRISPR locus (inactive CRISPR nuclease and array) and transposase protein are located within the left and right ends of the predicted transposon (e.g., FIG. 4B). In this case, effectors were predicted to guide DNA integration to specific genomic locations based on guide RNAs.

[0219] In some systems, effectors were predicted to have homology to documented CRISPR effector proteins but be inactive based on the absence of endonuclease domains and / or catalytic residues (e.g., FIG. 5, panel A). Transposases were predicted to associate with effectors when the CRISPR locus (inactive CRISPR nucleases and arrays) and the transposase protein was located within the left and right ends of the predicted transposon (FIG. 5, panels A and B).

[0220] Example 12 - CAST Identification CRISPR-associated transposons (CASTs) are a transposon-containing system that has evolved to interact with the CRISPR system to promote targeted integration of DNA cargo.

[0221] CAST is a genomic sequence that encodes one or more protein sequences involved in DNA transposition within the characteristic left and right ends of the transposon. Tn7 transposons as defined herein may contain the catalytic transposase TnsB, but may also contain the catalytic transposase TnsA, the loader proteins TnsC or TniB, and the target recognition proteins TnsD, TnsE, TniQ, and / or other transposon-associated components. The transposon ends contain predicted transposase binding sites, including direct and / or inverted repeats of 15 bp to 150 bp in length, flanking the transposon machinery and other "cargo" genes.

[0222] In addition, CAST also encodes DNA and / or RNA targeting CRISPR nuclease or effector near CRISPR array. In some systems, effector is predicted to be active nuclease based on the presence of endonuclease-associated catalytic domain and / or catalytic residue. In some systems, effector has sequence similarity with documented CRISPR effector protein, but is predicted to be inactive based on the absence of endonuclease domain and / or catalytic residue. Transposon is predicted to be associated with effector when CRISPR locus and transposon-associated protein are located within the left and right ends of predicted transposon. In this case, effector is predicted to guide DNA integration to specific genomic location based on guide RNA.

[0223] Example 13 - Class 2 Cas12K CAST The Cas12k CAST system encodes a nuclease-deficient CRISPR Cas12k effector, a CRISPR array, tracrRNA, and a Tn7-like transposition protein. Cas12k effectors are phylogenetically diverse, and features that confirm their association with CAST have been identified for some (Figure 8). For example, the left end of the transposon was identified downstream from the MG64-3 CRISPR locus, as indicated by a terminal inverted repeat and a self-matching spacer sequence (Figure 11A). The Cas12k CAST CRISPR repeat (crRNA) contains the conserved motif 5'-GNNGGNNTGAAAG-3' (Figure 9). Short repeat-antirepeat (RAR) motifs within the crRNA aligned with distinct regions of the tracrRNA (Figure 9 and Figures ​Figures10A-10B), and the RAR motifs appeared to define the start and end of the tracrRNA (e.g., for MG64-1, the 5' end of the tracrRNA contained RAR1 (TTTC) and the 3' end contained RAR2 (CCNNC) (Figure 10A).

[0224] Example 14 - Transposon end prediction Transposon ends were inferred from the intergenic regions flanking the effector and transposon machinery. For example, for Cas12k CAST, the intergenic regions located directly upstream from TnsB and directly downstream from the CRISPR locus were predicted to contain the left and right ends (LE and RE) of the Tn7 transposon.

[0225] Direct and inverted repeats (DR / IR) of approximately 12 bp were predicted on the contigs with up to two mismatches. In addition, short (approximately 10-20 bp) DR / IRs flanking the CAST transposon were found using the Dotplot algorithm. Matching DR / IRs located in intergenic regions flanking the CAST effector and transposon genes are predicted to encode transposon binding sites. LEs and REs extracted from the intergenic regions, encoding putative transposon binding sites, were aligned to define transposon end boundaries. Putative transposon LE and RE ends are: a) regions located within 400 bp upstream and downstream from the first and last predicted transposon-encoded genes, b) regions sharing multiple short inverted repeats, and c) regions sharing >65% nucleotide identity.

[0226] Example 15 - Single guide design Analysis of the intergenic regions surrounding the Cas effectors and CRISPR arrays identified potential anti-repeat sequences and a conserved "CYCC(n6)GGRG" stem-loop structure adjacent to the anti-repeat sequence corresponding to the double-stranded sequence of tracrRNA (Figure 11B). TracrRNA and crRNA repeats were folded and trimmed, and a tetraloop sequence of GAAA was added to maintain the stem-loop region of the crRNA-tracrRNA complementary sequence.

[0227] Example 16 - In vitro integration activity using targeted nucleases In situ expression and protein sequence analysis showed that several RNA-guided effectors were active nucleases. They contained predicted endonuclease-associated domains (matching the RuvC and HNH endonuclease domains) and / or predicted HNH and RuvC catalytic residues. Candidate activity was tested using an in vitro expression system and in vitro transcribed RNA with engineered single guide RNA sequences. Active proteins that successfully cleaved the library gave rise to a band of approximately 170 bp in the gel.

[0228] Example 17 - Programmable DNA integration CAST activity was tested using five types of components: (1) Cas effector protein expressed by an in vitro expression system (SEQ ID NO:1), (2) target DNA fragment or plasmid (SEQ ID NO:31) containing the target sequence and PAM corresponding to the Cas enzyme, (3) donor DNA fragment (SEQ ID NO:8-11) containing markers or fragments of DNA flanking the LE and RE of the transposase system in the DNA fragment or plasmid, (4) any combination of transposase proteins expressed using an in vitro expression system (SEQ ID NO:2-4), and (5) an engineered in vitro transcribed single guide RNA sequence (SEQ ID NO:5). Successful transposition of the donor fragment was assayed for by PCR amplification of the donor-target junction.

[0229] After performing the transposition reaction, PCR amplification of the junctions indicated that proper donor-target formation had occurred and that the transposition reaction was sg-dependent (Figure 9). PCR amplification of reactions #3 and #4 indicated that both orientations of the donor to the target had been achieved, i.e., LE closer to the PAM and RE closer to the PAM. Although both transposition orientations occurred, there appeared to be a preference for donor integration in the target with the LE closer to the PAM, as represented by the strong bands present for reactions #4 and #5.

[0230] Sanger sequencing of the preferred orientation products was performed. Of the integrations that occurred with the LE closer to the PAM, there was a clear degradation of the sequencing chromatogram signal from either the forward or reverse orientation across the target / donor junction. This indicated that of the products oriented with the LE closer to the PAM, integration occurred over a range of nucleotides, with the major product of the LE closer to the PAM as a 61-bp integration from the PAM (Figure 10A). Sequencing of the donor product across the donor-target junction defined the composition of the essential outer border of the LE and RE sequences (Figure 10A,B). Sequencing of the RE on the product of the LE closer to the PAM showed a 3-bp overlap downstream of the donor RE (Figure 10B). This is due, in part, to a Tn7 transposase integration event that cut and ligated the donor fragment at the staggered cut sites. The 3-bp overlap is smaller than the expected 5-bp overlap from other Tn7 transposases.

[0231] Sanger sequencing of PCR amplification products against the 8N library of target plasmids also demonstrated the PAM preference of the MG64-1 effector as nGTn / nGTt on the 5' end of the spacer (Figure S10C). NGS analysis of the PAM library targets confirmed the nGTn motif preference at the 5' end.

[0232] Further development of single-guide studies confirmed the activity of MG64-1 with the new sgRNA scaffold (Figures 13A-13C).

[0233] Example 18 - Determining the Integration Window The PCR junctions of the amplified PAM were indexed and sequenced for the NGS libraries. Reads were mapped and quantified using CRISPResso using the amplicon sequence of the putative transposition sequence with an integrated distance of 60 bp from the PAM (guideseq=20 bp 3' end of LE or RE, window center=0, window size=20). Indel histograms were normalized to the total indel reads detected and the frequency was plotted against the 60 bp reference sequence (Figure 14).

[0234] Both PCR reaction 5 (LE proximal to the PAM, top panel of FIG. 14) and PCR 4 (RE distal to the PAM, bottom panel of FIG. 14) were plotted over sequence and distance from the PAM for MG64-1. Analysis of the integration window indicates that 95% of integrations that occurred at the spacer PAM site were within a 10 bp window 58-68 nucleotides away from the PAM. The difference in integration distance between distal and proximal frequencies reflected overlap of integration sites, i.e., an overlap of 3-5 base pairs as a result of the staggered nuclease activity of the transposase during integration.

[0235] Example 19 - Colony PCR screening for transposase activity Transposition activity was assayed via colony PCR screening. After transformation with the pDonor plasmid, E. coli were plated on LB-agar containing ampicillin, chloramphenicol, and tetracycline. Picked CFU were added to a solution containing PCR reagents and primers flanking the selected insertion junction. PCR reactions of integration products were visible on a gel (Figure 15). Sequencing results of picked colony PCR products confirmed that they represented transposition events, as they spanned the junction between the LE and PAM at the engineered target site within the lacZ gene (Figure 16).

[0236] Example 20 - Single guide operation The predicted RNA folding of active single RNA sequences was calculated at 37° using the method of Andronescu 2007. All hairpin loop secondary structures were singly removed from the construct and compiled iteratively into smaller single guides. Engineered single guides (esg) 4, 6, 7, 8, 9 were active for donor transposition (panels C and D of FIG. 17), and engineered sgRNAs 8 and 9 were weaker single guides and transposed with PCR5 (panel D of FIG. 17). Engineered guide 5 was able to transpose, but engineered sgRNA 10 transposed weakly with PCR5 (FIG. 17). Esg17 is a combination of deletions in esg6 and esg7, and esg18 is a combination of esg4 and esg5. Both were able to transpose strongly across both PCRs 4 and 5 (Figure 17, panels G and H), but the combined addition of esg6 and esg18 to create esg19 resulted in weaker transposition in PCR 5, and the addition of esg7 to esg19 to create esg20 resulted in a very weak junction of transposition for PCR 5 (Figure 17, panels G and H). In a second approach, the tracrRNA of MG64-1 was aligned to the documented Vk-type tracrRNA and the region of the unique insertion was mutated from the single guide. The sgRNA was minimized by truncation of the insertion sequence of the MG64-1 sgRNA (Figure 14). Two subsequent deletions, esg2 and esg3, were also tested (Figure 17, panels A and B), but neither esg2 nor esg3 resulted in significant transposition, and thus the single guide was minimized by 57 bases.

[0237] Example 21 - LE-RE Minimization Sequencing of the target-transposition junction assisted in the identification of terminal inverted repeats by identifying the outermost sequence from the donor plasmid incorporated in the targeting reaction. By performing a 14 bp repeat analysis with 10% variability, short repeats contained within the termini were identified and truncations of these minimal termini were designed to retain the repeats while deleting the extra sequence. Multiple rounds of prediction and cloning were performed and each interaction was tested by in vitro transposition. Initial LE and RE deletions were designed independently and cloned at 68 bp, 86 bp, and 105 bp for the LE and 178 bp, 196 bp, and 242 bp for the RE. The RE of 64-1 also had a significant stretch of sequence that was devoid of repeats, so both 50 bp and 81 bp internal deletions were designed and cloned. Transposition between all single deletions was robust for both PCR4 and PCR5 (panels A and B of FIG. 18), after which an 81 bp internal deletion was pursued with a combined deletion of the RE. The previous trimmed ends of 178, 196, and 212 bp were cloned onto the 81 bp internal deletion and transposition was tested. Transposition was active for all constructs designed. In combination with a 68 bp LE, transposition proved to be active up to a 68 bp LE region combined with a 96 bp RE region (panels E and F of FIG. 18).

[0238] Example 22 - Overhang effect of dislocation To test whether extra sequences outside the TnsB binding motif are required for transposition, oligos designed for the TGTACA motif in both the LE and RE were designed and synthesized with 0, 1, 2, 3, 5, and 10 bp of extra base pairs. These synthetic oligos were used to generate donor PCR fragments with overhangs and tested for their ability to transpose into the target site. Most notably, PCR6 was rarely detected from in vitro reactions (Figure 18, panel G, lanes 1, 2), but efficient integration with PCR6 was detected with small 0-3 bp overhangs, reflecting the orientation of the RE proximal to the PAM that is not detected with larger flanking sequences.

[0239] Example 23 - CAST NLS Design Eukaryotic genome editing for therapeutic purposes relies primarily on the import of editing enzymes into the nucleus. Small polypeptide extensions of larger proteins signal to cellular components for protein import across the nuclear membrane. The placement of these tags is not trivial as import function versus function of the protein to which it is fused is a potential trade-off depending on the location of the NLS tag. To test the functional orientation of the NLS to each of the components of the CAST complex, constructs were designed and synthesized in which a nucleoplasmic NLS was fused to the N-terminus and an SV40 NLS was fused to the C-terminus of each of the components of MG CAST. Proteins from these constructs were expressed in cell-free in vitro transcription / translation reactions and tested for in vitro transposition activity with a complementary set of untagged components. NLS-tagged constructs were assessed for maintenance of activity by PCR of donor-target junctions using PCR4 (to assess RE-distal transposition) and the cognate transposition event PCR5 (LE-to-proximal transposition).

[0240] Most components yielded a single NLS orientation that maintained activity. TnsB was a CAST component that was active with both N- and C-terminal NLSs by both PCR4 and PCR5 (Figure 19, panels A and B). TniQ was active with an N-terminal NLS tag (Figure 19, panels C and D). And the Cas12k component was active with a C-terminal tagged NLS (Figure 19, panels E and F, lanes 5, 6). Further development of Cas12k containing both nucleoplasmic and SV40 NLS tags was tested and found to be active (Figure 19, panels I and J, lane 4). TnsC was weakly active with an N-terminal NLS (Figure 19, panels E and F, lane 7), but further exploration of TnsC tagging identified new functional NLS-HA-TnsC and NLS-FLAG-TnsC constructs (Figure 19, panels G and H, lanes 3 and 7, respectively). The end result was a set of complete NLS-tagged components that were active in vitro in both NLS-TnsB and TnsB-NLS orientations (FIG. 20, lanes 5, 6, panels A and B).

[0241] Example 24 - Design and testing of Cas12k and TniQ protein fusion constructs In an attempt to simplify the expression of protein components and minimize the delivery of these components into cells, fusion constructs were designed, synthesized, and tested between the Cas12k effector and the TniQ protein. Both orientations of TniQ fused to Cas12k were designed and a C-terminal fusion, Cas-TniQ, and an N-terminal fusion, TniQ-Cas, were synthesized. Both constructs were weakly active with PCR4 (panel A of FIG. 21), but when expressed in vitro and assayed for transposition ability, the PCR5 junction was robustly formed by the TniQ-Cas fusion protein (panel B of FIG. 21). Transposition length was assayed with variable linker domains including original (20 amino acid linker), 48, 68, 72, and 77 (panels C, D, E, and F of FIG. 21). An NLS tag was then ligated to the N-terminus of TniQ and the C-terminus of Cas12k and found to still be active by PCR5 (panels E and F of FIG. 21).

[0242] Two other linkers were used to fuse the effector gene and the TniQ gene: the self-terminating translation sequence, P2A, was active in the Cas-NLS-P2A-NLS-TniQ construct (Figure 21, panels G and H, lane 6), and the MCV internal ribosome entry sequence (IRES) mRNA-based linker allowed independent translation of the two components in cells (Figure 23, panels F and G).

[0243] Example 25 - Intracellular expression-coupled in vitro translocation assay To test the functionality of the NLS constructs in a physiologically relevant environment, constructs cloned with active NLS-tagged CAST components were integrated into K562 cells using lentiviral transduction. Briefly, constructs cloned into lentiviral transfer plasmids were transfected into 293T cells with envelope and packaging plasmids, and virus containing supernatants were harvested from the medium after 72 hours of incubation. The virus-containing medium was then incubated with K562 cell lines containing 8 μg / mL polybrene for 72 hours, and transfected cells were then selected for co-integration using puromycin at 1 μg / mL for 4 days. The selected cell lines were harvested at the end of the 4 days and differentially lysed for nuclear and cytoplasmic fractions. Subsequent fractions were then tested for transposition competence using a complementary set of in vitro expression components.

[0244] Both NLS-TnsB and TnsB-NLS were tested by cell fractionation and in vitro translocation, and translocation was detected across both cytoplasmic and nuclear fractions, with NLS-TniQ having detectable activity in the cytoplasm (Figure 22, Panels A and B). NLS-HA-TnsC and NLS-FLAG-TnsC were active in both cytoplasmic and nuclear fractions when expressed (Figure 22, Panel D), whereas PCR4 was formed in the nuclear fraction for both TnsC constructs (Figure 22, Panel C).

[0245] When both NLS-TnsB or TnsB-NLS were linked to NLS-FLAG-TnsC by using an IRES, NLS-TnsB-IRES-NLS-FLAG-TnsC was mostly active in the nuclear fraction, while TnsB-NLS-IRES-NLS-FLAG-TnsC was active in both the cytoplasmic and nuclear fractions, indicating that NLS-TnsB has a higher ability to transport to the nucleus (Figure 21, panels E and F).

[0246] Cas12k fusions in cells were similarly fractionated and tested for translocation. Cas-NLS Cas-NLS-P2A-NLS-TniQ was transduced into cells, fractionated, and tested in vitro for intracellular activity. Cas-NLS-P2A-NLS-TniQ could be translocated into the cytoplasm by adding a single guide to the reaction (Panel A of FIG. 23). The Cas-NLS-P2A-NLS-TniQ construct in the nuclear fraction was complemented by supplementing with holo-Cas protein (+sgRNA) or additional TniQ with sgRNA. This indicates that both Cas-NLS and NLS-TniQ are making their way to the nucleus (Panels B and C of FIG. 23). The NLS-TniQ-Cas-NLS fusion protein had similar results but required additional supplementation with TniQ (Figure 23, panels D and E), and Cas-NLS-IRES-NLS-TniQ required supplementation from holo-Cas-NLS only (Figure 23, panels F and G). Overall, this indicates that all components of CAST could be delivered to the nuclear fraction of cells.

[0247] Example 26 - Transposon end verification via gel shift To verify the activity of TnsB on the predicted transposon end sequences, FAM-labeled oligos were used to amplify the LE of MG64-1. Using a cell-free transcription / translation system, MG64-1 TnsB protein was expressed and incubated with the LE FAM-labeled product. After 30 min of incubation, binding was observed on a native 5% TBE gel (Figure 24). Multiple bands of fluorescent product in the co-incubated lanes (Figure 24, lane 3) indicated a minimum of two TnsB binding sites.

[0248] The disclosed system can be used for a variety of applications, such as, for example, nucleic acid editing (e.g., gene editing) or binding (e.g., sequence-specific binding) to nucleic acid molecules. Such systems can be used, for example, to correct (e.g., remove or replace) genetically inherited mutations that may cause disease in a subject, to inactivate genes to confirm their function in cells, as diagnostic tools to detect disease-causing genetic elements (e.g., via cleavage of reverse-transcribed viral RNA or amplified DNA sequences encoding disease-causing mutations), as inactivated enzymes combined with probes to target and detect specific nucleotide sequences (e.g., sequences encoding antibiotic resistance in bacteria), to inactivate viruses by targeting viral genomes or to render them unable to infect host cells, to add genes or modify metabolic pathways to engineer organisms to produce valuable small molecules, macromolecules, or secondary metabolites, to establish gene drive elements for evolutionary selection, and / or to detect cellular perturbations by exogenous small molecules and nucleotides as biosensors.

[0249] Example 27 - Class 2 Cas12k CAST-based predictions The Cas12k CAST system encodes a nuclease-deficient CRISPR Cas12k effector, a CRISPR array, tracrRNA, and a Tn5053-like transposition protein (Figure 25A). Cas12k effectors are phylogenetically diverse and have been characterized that establish their association with CAST (Figures 25A-25B). For example, the left end of the transposon was identified downstream from many Cas12k effectors and their CRISPR loci, as indicated by terminal inverted repeats and self-matching spacer sequences (Figures 25A-25B).

[0250] The transposon ends of the Cas12k CAST system were determined from the intergenic regions adjacent to the CRISPR locus and the transposon machinery. For example, the intergenic regions located directly upstream from TnsB and directly downstream from the CRISPR locus were predicted to contain the left and right ends (LE and RE) of the transposon. These intergenic regions were aligned across several homologs and the conserved regions were used to predict the transposon end boundaries (Figures 26A-26B).

[0251] The 3' end of the Cas12k CAST CRISPR repeat (crRNA) contains a conserved motif 5'-GNNGGNNTGAAAG-3' when aligned between homologs, which are predicted to bind to different regions of the tracrRNA to form secondary and tertiary guide RNA structures (Figures 27 and 28). Self-matching spacers within the CAST transposon are often found next to pseudo-CRISPR repeats near the CRISPR array (Figure 25A, bottom alignment).

[0252] Analysis of the intergenic regions surrounding the Cas effectors and CRISPR arrays identified potential anti-repeat sequences and a conserved "CCYCC(n6)GGRGG" stem-loop structure adjacent to the anti-repeat sequence corresponding to the double-stranded sequence of tracrRNA (Figure 27). A covariance model was constructed using good quality tracrRNA and searched in all Cas12k CAST genomic fragments identified in this study.

[0253] For single guide RNAs (sgRNAs), the tracrRNA and crRNA repeats were folded and trimmed, and a tetraloop sequence of GAAA was added to maintain the stem-loop region of the crRNA-tracrRNA complementary sequence (Figure 28). In general, sgRNAs share conserved structural features despite sharing less than 70% pairwise nucleotide identity (Figure 27).

[0254] Example 28 - In vitro characterization of the Cas12k CAST system To test the functionality of the Cas12k CAST system and elucidate potential PAMs, transposition reactions were assembled using synthetic Cas12k effectors and Tn5053-like proteins under the control of a T7 promoter. Each open reading frame was expressed in vitro using an in vitro expression system and assembled in a transposition reaction using transposition buffer, donor PCR fragments, and a plasmid-based target with the 8N target library (Figure 29A). When the CAST system is active and able to transpose the donor fragment into the library of target plasmids, the transposition reaction can be PCR amplified to recover each donor-target junction of the two potential products of transposition (Figure 29B).

[0255] Of the Cas12k CAST systems discovered, three were prioritized for their novelty and completeness and tested for transposition potential in vitro. MG64-1 CAST was able to transpose cargo into the donor plasmid in an sgRNA-dependent manner (Figure 30A). For each potential junction PCR amplification, it was predicted that all four junctions would be observed if both orientations of integration were complete, or only two of the junctions would be observed if integration only occurred in a single orientation (Figure 29B). Surprisingly, robust transposition was observed in three of the four junction PCR reactions. The observed reactions represent both potential left-end junction products in both target-LE-cargo-RE (T-LR) and target-RE-cargo-LE (T-RL) orientations (PCR5 and PCR3, respectively), as well as the T-LR-oriented right-end product (PCR4) (Figure 30A).

[0256] Sanger-based sequencing of the PCR transposition fragments aligned precisely to the target sequence, and the sequencing signal degraded rapidly at the target-donor junction (Figure 41). Signal degradation indicated a population of integration products, and sequencing from the donor end of the transposition products verified both the LE and RE predictions in the MG64-1 system.

[0257] To elucidate the PAM preference of the active CAST system, successful transposition events from a population of 8N randomized libraries that received donor sequences were sequenced via NGS. Sequencing reads identified the GTN-5' PAM ...

Claims

1. 1. A system for translocating a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, comprising: a) a Cas effector complex comprising a class 2 V-type Cas effector, a small prokaryotic ribosomal protein subunit S15, and an engineered guide polynucleotide configured to hybridize to the target nucleic acid site; b) a Tn7-type transposase complex configured to bind to the Cas effector complex, the Tn7-type transposase complex comprising TnsB, TnsC, and TniQ components; c) a double-stranded nucleic acid configured to interact with the Tn7-type transposase complex and comprising the cargo nucleotide sequence.

2. the Cas effector complex (a) non-covalently bound to the Tn7-type transposase complex; (b) covalently linked to the Tn7-type transposase complex or fused to the Tn7-type transposase complex.

3. A system described in any one of claims 1 or 2, wherein the cargo nucleotide sequence is adjacent to a left transposase recognition sequence and a right transposase recognition sequence recognized by the Tn7-type transposase complex.

4. The system described in claim 3, wherein the left recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78, and the right recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93.

5. The system described in claim 1, wherein the target nucleic acid comprises a PAM sequence compatible with the Cas effector complex.

6. The system described in claim 5, wherein the PAM sequence comprises sequence number 31.

7. A system described in claim 5 or 6, wherein the PAM sequence is located approximately 50 to approximately 70 base pairs from the target nucleic acid site.

8. The system described in claim 7, wherein the PAM sequence is located 3' of the target nucleic acid site, or the PAM sequence is located 5' of the target nucleic acid site.

9. The system described in claim 1, wherein the class 2 V-type Cas effector comprises a polypeptide having an array having at least 80% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220.

10. The system described in claim 1, wherein the class 2 V-type Cas effector comprises a polypeptide having any one of the sequences of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220.

11. The system described in claim 1, wherein the TnsB component comprises a polypeptide having an sequence having at least 80% identity to any one of SEQ ID NOs: 2, 13, 17, and 65.

12. The system described in claim 1, wherein the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least 80% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.

13. The system described in claim 1, wherein the engineered guide polynucleotide comprises a sequence comprising at least about 46 to 80 consecutive nucleotides having at least 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222.

14. The system described in claim 1, wherein the engineered guide polynucleotide comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, and 123-140.

15. The system described in claim 1, wherein the small prokaryotic ribosomal protein subunit S15 comprises a sequence having at least 80% sequence identity with any one of SEQ ID NOs: 181-183 or 187-189.

16. The system described in claim 1, wherein the class 2 V-type Cas effector and the Tn7-type transposase complex are encoded by a polynucleotide sequence comprising less than about 10 kilobases.

17. A method for translocating a cargo nucleotide sequence within a target nucleic acid site, comprising introducing the system described in claim 1 into a cell.

18. A cell comprising the system described in claim 1.

19. The cell described in claim 18, wherein the cell is a eukaryotic cell, a mammalian cell, an immortalized cell, an insect cell, a yeast cell, a plant cell, a fungal cell, or a prokaryotic cell.

20. The cell of claim 19, wherein the cell is A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, Saos, C2C12, L cells, HT1080, HepG2, Huh7, K562, primary cells, or derivatives thereof.