Systems and methods for rearranging cargo nucleotide sequences

The use of a Tn7-type transposase complex and Cas effector system with engineered guide polynucleotides addresses inefficiencies in nucleotide sequence manipulation, enabling precise transposition of cargo sequences to target sites.

JP7863897B2Active Publication Date: 2026-05-22METAGENOMI THERAPEUTICS INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
METAGENOMI THERAPEUTICS INC
Filing Date
2021-08-23
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Current methods for manipulating nucleotide sequences using CRISPR/Cas systems are limited in efficiency and specificity, particularly in the transposition of cargo nucleotide sequences to target sites.

Method used

A system comprising a Tn7-type transposase complex, a Cas effector complex, and a manipulated guide polynucleotide is used to facilitate the transposition of cargo nucleotide sequences to target nucleic acid sites, utilizing a class II, type V Cas effector and engineered guide polynucleotides for precise hybridization and binding.

Benefits of technology

The system achieves efficient and specific transposition of cargo nucleotide sequences to target nucleic acid sites, enhancing the precision and effectiveness of nucleotide manipulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007863897000001
    Figure 0007863897000001
  • Figure 0007863897000002
    Figure 0007863897000002
  • Figure 0007863897000003
    Figure 0007863897000003
Patent Text Reader

Abstract

The present disclosure provides systems and methods for translocating a cargo nucleotide sequence to a target nucleic acid site. These systems and methods may include a first double-stranded nucleic acid comprising a cargo nucleotide sequence configured to interact with a recombinase complex, a cas effector complex comprising a cas effector and at least one engineered guide polynucleotide configured to hybridize to the target nucleic acid site, and a recombinase complex configured to recruit the cargo nucleotide to the target nucleic acid site.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 082,983, entitled "SYSTEMS AND METHODS FOR TRANSPOSING CARGO NUCLEOTIDE SEQUENCES," filed on September 24, 2020; U.S. Provisional Patent Application No. 63 / 187,290, entitled "SYSTEMS AND METHODS FOR TRANSPOSING CARGO NUCLEOTIDE SEQUENCES," filed on May 11, 2021; and U.S. Provisional Patent Application No. 63 / 232,578, entitled "SYSTEMS AND METHODS FOR TRANSPOSING CARGO NUCLEOTIDE SEQUENCES," filed on August 12, 2021, each of which is hereby incorporated by reference in its entirety.

Background Art

[0002] Cas enzymes, along with their associated Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) guide ribonucleic acid (RNA), are thought to be a widespread (about 45% of bacteria, about 84% of archaea) component of the prokaryotic immune system and play a role in protecting the microorganism against non-self nucleic acids such as infecting viruses and plasmids by CRISPR-RNA-guided nucleic acid cleavage. Deoxyribonucleic acid (DNA) elements encoding CRISPR RNA elements can be relatively conserved in structure and length, while their CRISPR-associated (Cas) proteins are highly diverse and contain various nucleic acid interaction domains. Although CRISPR DNA elements were discovered in 1987, the programmable endonuclease cleavage ability of the CRISPR / Cas complex has only recently been recognized, and recombinant CRISPR / Cas systems are being used in a variety of DNA manipulation and gene editing applications.

[0003] Sequence List This application includes a sequence listing filed electronically in ASCII format, which is incorporated herein by reference in its entirety. The above ASCII copy, created on 20 August 2021, is named 55921-714_602_SL.txt and is 196,492 bytes in size. [Overview of the project]

[0004] In some embodiments, the Disclosure provides a system for transposing a cargo nucleotide sequence to a target nucleic acid site, the system comprising: a first double-stranded nucleic acid comprising a cargo nucleotide sequence configured to interact with a Tn7-type transposase complex; a Cas effector complex comprising a class II, type V Cas effector and an engineered guide polynucleotide configured to hybridize to the target nucleotide sequence; and a Tn7-type transposase complex configured to bind to the Cas effector complex, the Tn7-type transposase complex comprising a TnsB subunit. In some embodiments, the cargo nucleotide sequence is adjacent to a left transposase recognition sequence and a right transposase recognition sequence. In some embodiments, the system further comprises a second double-stranded nucleic acid comprising the target nucleic acid site. In some embodiments, the system further comprises a PAM sequence that fits the Cas effector complex adjacent to the target nucleic acid site. In some embodiments, the PAM sequence is located at 3' of the target nucleic acid site. In some embodiments, the PAM sequence is located at 5' of the target nucleic acid site. In some embodiments, the manipulated guide polynucleotide is configured to bind to the class II, type V Cas effector. In some embodiments, the class II, type V Cas effector comprises a polypeptide containing a sequence having at least 80% identity to SEQ ID NOs. 1, 12, 16, 20-30, 64, or 80-85 or their variants. In some embodiments, the TnsB subunit comprises a polypeptide containing a sequence having at least 80% identity to SEQ ID NOs. 2, 13, 17, or 65 or their variants. In some embodiments, the Tn7 type transposase complex comprises at least one, at least two, or three polypeptides containing a sequence having at least 80% identity to any one of SEQ ID NOs. 3-4, 14-15, 18-19, or 66-67 or their variants.In some embodiments, the manipulated guide polynucleotide includes a sequence comprising at least 46 to 80 consecutive nucleotides having at least 80% identity to one of sequence numbers 5-6, 32-33, 94-95, or 104-105 or its variants. In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 80% sequence identity to one of sequence numbers 106, 107, 108, 5, 45-63, 68-75, or 96-103 or its variants (non-degenerate nucleotides). In some embodiments, the left recombinase sequence includes a sequence having at least 80% identity to sequence numbers 9, 11, 36-38, 76, or 78 or its variants. In some embodiments, the right recombinase sequence includes a sequence having at least 80% identity to sequence numbers 8, 10, 39-44, 77, 79, or 93 or its variants. In some embodiments, the class II, type V Cas effector and the Tn7 type transposase complex are encoded by a polynucleotide sequence containing less than approximately 10 kilobases.

[0005] In some embodiments, the Disclosure provides a method for transposing a cargo nucleotide sequence to a target nucleic acid site containing a target nucleotide sequence, the method comprising the steps of expressing any of the embodiments or systems described herein in a cell, or introducing any of the embodiments or systems described herein into a cell.

[0006] In some embodiments, the Disclosure discloses a method for transposing a cargo nucleotide sequence to a target nucleic acid site, the method comprising contacting a first double-stranded nucleic acid containing the cargo nucleotide sequence with a Cas effector complex comprising a Cas effector complex comprising a class II, type V Cas effector and at least one manipulated guide polynucleotide configured to hybridize to the target nucleotide sequence, and a Tn7 type transposase complex configured to bind to the Cas effector complex, comprising a TnsB subunit, and the second double-stranded nucleic acid containing the target nucleic acid site. In some embodiments, the cargo nucleotide sequence is adjacent to a left transposase recognition sequence and a right transposase recognition sequence. In some embodiments, the system further includes a PAM sequence that fits the Cas effector complex adjacent to the target nucleic acid site. In some embodiments, the PAM sequence is located at 3' of the target nucleic acid site. In some embodiments, the manipulated guide polynucleotide is configured to bind to the class II, type V Cas effector. In some embodiments, the class II, type V Cas effector comprises a polypeptide having at least 80% identity to sequence numbers 1, 12, 16, 20-30, 64, or 80-85, or their variants. In some embodiments, the TnsB subunit comprises a polypeptide having at least 80% identity to sequence numbers 2, 13, 17, or 65, or their variants. In some embodiments, the Tn7 type transposase complex comprises at least one or at least two polypeptides having at least 80% identity to any one of sequence numbers 3-4, 14-15, 18-19, or 66-67, or their variants.In some embodiments, the manipulated guide polynucleotide comprises a sequence containing at least 46 to 80 consecutive nucleotides having at least 80% identity to any one of SEQ ID NOs. 5-6, 32-33, 94-95, or 104-105 or its variants. In some embodiments, the left-hand recombinase sequence comprises a sequence having at least 80% identity to SEQ ID NOs. 9, 11, 36-38, 76, or 78, or its variants. In some embodiments, the right-hand recombinase sequence comprises a sequence having at least 80% identity to SEQ ID NOs. 8, 10, 39-44, 77, 79, or 93, or its variants. In some embodiments, the class II, type V Cas effector and the Tn7 type transposase complex are encoded by a polynucleotide sequence containing less than about 10 kilobases.

[0007] In some embodiments, the Disclosure provides a system for transposing a cargo nucleotide sequence to a target nucleic acid site, the system comprising: a first double-stranded nucleic acid comprising a cargo nucleotide sequence configured to interact with a Tn7 type transposase complex; a Cas effector complex comprising a class II, type V Cas effector and an engineered guide polynucleotide configured to hybridize to the target nucleotide sequence; and a Tn7 type transposase complex configured to bind to the Cas effector complex, wherein TnsB, TnsC, and Tni The transposase complex comprises a Tn7 type transposase complex containing a Q component, wherein (a) the class II, type V Cas effector comprises a polypeptide having a sequence having at least 80% sequence identity to one of SEQ ID NOs: 1, 12, 16, 20-30, 64, or 80-85 or a variant thereof, or (b) the Tn7 type transposase complex comprises a TnsB, TnsC, or TniQ component having a sequence having at least 80% sequence identity to one of SEQ ID NOs: 2-4, 13-15, 17-19, or 65-67 or a variant thereof. In some embodiments, the transposase complex is non-covalently bonded to the Cas effector complex. In some embodiments, the transposase complex is covalently bonded to the Cas effector complex. In some embodiments, the transposase complex is fused to the Cas effector complex in a single polypeptide. In some embodiments, the class II, type V Cas effector comprises a polypeptide having a sequence having at least 80% sequence identity with one of sequence numbers 1, 12, 16, 20-30, 64, or 80-85 or a variant thereof. In some embodiments, the Tn7 type transposase complex comprises a TnsB, TnsC, or TniQ component having a sequence having at least 80% sequence identity with one of sequence numbers 2-4, 13-15, 17-19, or 65-67 or a variant thereof. In some embodiments, the class II, type V Cas effector is a Cas12k effector.In some embodiments, the cargo nucleotide sequence is adjacent to a left transposase recognition sequence and a right transposase recognition sequence. In some embodiments, the system further comprises a second double-stranded nucleic acid containing the target nucleic acid site. In some embodiments, the system further comprises a PAM sequence that fits the Cas effector complex adjacent to the target nucleic acid site. In some embodiments, the PAM sequence is located at 5' or 3' of the target nucleic acid site. In some embodiments, the PAM sequence includes SEQ ID NO: 31. In some embodiments, the manipulated guide polynucleotide is configured to bind to the class II, type V Cas effector. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least about 46 to 80 consecutive nucleotides having at least 80% identity to one of SEQ ID NOs: 5-6, 32-33, 94-95, or 104-105 or their variants. In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 80% sequence identity to one of the non-degenerate nucleotides of SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, or 96-103 or their variants. In some embodiments, the left-hand recombinase sequence includes a sequence having at least 80% identity to one of the SEQ ID NOs: 9, 11, 36-38, 76, or 78 or their variants. In some embodiments, the right-hand recombinase sequence includes a sequence having at least 80% identity to one of the SEQ ID NOs: 8, 10, 39-44, 77, 79, or 93. In some embodiments, the class II, type V Cas effector and the Tn7 type transposase complex are encoded by a polynucleotide sequence containing less than approximately 10 kilobases.In some embodiments, (a) the Class II, Type V Cas effector comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1, 81, 82, 83, or 85 or a variant thereof; (b) the left recombinase sequence comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 9, 11, 36, 37, or 38 or a variant thereof; (c) the right recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 8, 39, 40, 41, 42, 43, 44, or 93 or a variant thereof; d) The manipulated guide polynucleotide comprises (i) a sequence having at least 80% sequence identity to at least about 46 to 80 nucleotides of SEQ ID NO: 6 or its variants, or (ii) a sequence having at least 80% identity to any one of SEQ ID NOs: 5, 45 to 63, 68 to 75, or 96 to 103 or its variants, or (e) the TnsB, TnsC, and TniQ components comprising polypeptides having sequences having at least 80% identity to SEQ ID NOs: 2 to 4 or their variants, or (f) the PAM sequence comprising SEQ ID NO: 31.In some embodiments, (a) the class II, type V Cas effector comprises a sequence having at least 80% sequence identity to SEQ ID NO: 12 or its variants; (b) the left recombinase sequence comprises a sequence having at least 80% sequence identity to SEQ ID NO: 76 or its variants; (c) the right recombinase sequence comprises a sequence having at least 80% identity to SEQ ID NO: 77 or its variants; (d) the manipulated guide polynucleotide comprises (i) a sequence having at least 80% sequence identity to at least about 46 to 80 nucleotides of SEQ ID NO: 32 or 104 or its variants; or (ii) a sequence having at least 80% identity to the non-degenerate nucleotide of either SEQ ID NO: 107 or 102 or its variants; or (e) the TnsB, TnsC, and TniQ components comprises polypeptides having sequences having at least 80% identity to SEQ ID NO: 13 to 15 or its variants. In some embodiments, (a) the class II, type V Cas effector comprises a sequence having at least 80% sequence identity to SEQ ID NO: 16 or its variants; (b) the left recombinase sequence comprises a sequence having at least 80% sequence identity to SEQ ID NO: 78 or its variants; (c) the right recombinase sequence comprises a sequence having at least 80% identity to SEQ ID NO: 79 or its variants; (d) the manipulated guide polynucleotide comprises (i) a sequence having at least 80% sequence identity to at least about 46 to 80 nucleotides of SEQ ID NO: 33 or 105 or its variants; or (ii) a sequence having at least 80% identity to the non-degenerate nucleotide of either SEQ ID NO: 108 or 103 or its variants; or (e) the TnsB, TnsC, and TniQ components comprise polypeptides having sequences having at least 80% identity to SEQ ID NO: 17 to 19 or their variants.

[0008] In some embodiments, the Disclosure provides an engineered nuclease system, the engineered nuclease system comprising an endonuclease comprising a RuvC domain, wherein the endonuclease is a class II, type VK Cas effector derived from an uncultured microorganism and having at least 80% identity to one of sequence numbers 1, 12, 16, 20-30, 64, or 80-85 or its variants, and an engineered guide, the engineered guide RNA comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence. In some embodiments, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least 80% identity to one of sequence numbers 5-6, 32-33, 94-95, or 104-105 or its variants. In some embodiments, the manipulated guide polynucleotide comprises a sequence having at least 80% identity to one of the non-degenerate nucleotides of sequence numbers 106, 107, 108, 5, 45-63, 68-75, or 96-103 or their variants. In some embodiments, the system further comprises a PAM sequence that fits the Cas effector complex adjacent to the target nucleic acid site. In some embodiments, the PAM sequence is located 5' of the target nucleic acid site. In some embodiments, the PAM sequence comprises sequence number 31.In some embodiments, (a) the class II, type VK Cas effector comprises a sequence having at least 80% sequence identity with any one of sequence numbers 1, 81, 82, 83, or 85 or a variant thereof; (b) the left recombinase sequence comprises a sequence having at least 80% sequence identity with any one of sequence numbers 9, 11, 36, 37, or 38 or a variant thereof; (c) the right recombinase sequence comprises a sequence having at least 80% identity with any one of sequence numbers 8, 39, 40, 41, 42, 43, 44, or 93 or a variant thereof; d) The manipulated guide polynucleotide comprises (i) a sequence having at least 80% sequence identity to at least about 46 to 80 nucleotides of SEQ ID NO: 6 or its variants, or (ii) a sequence having at least 80% identity to any one of SEQ ID NOs: 5, 45 to 63, 68 to 75, or 96 to 103 or its variants, or (e) the TnsB, TnsC, and TniQ components comprising polypeptides having sequences having at least 80% identity to SEQ ID NOs: 2 to 4 or their variants, or (f) the PAM sequence comprising SEQ ID NO: 31.

[0009] Further aspects and advantages of this disclosure will be readily apparent to those skilled in the art from the detailed description below, and only exemplary embodiments of this disclosure are shown and described here. As will be understood, this disclosure may also be possible in other embodiments and different embodiments, and various details thereof may be modified in various obvious ways without all departing from this disclosure. Accordingly, the drawings and description are intended to be illustrative and not limiting.

[0010] Reference All publications, patents, and patent applications referenced herein are incorporated herein by reference to the same extent as each individual publication, patent, or patent application is specifically and individually incorporated herein by reference. [Brief explanation of the drawing]

[0011] Novel features of the present invention are expressed in particular in the appended claims. The features and advantages of the present invention will be better understood by referring to the following detailed description illustrating exemplary embodiments in which the principles of the present invention are used, and to the following appended drawings (also referred to herein as “Figure” and “FIG.”).

[0012] [Figure 1] This diagram illustrates the typical organization of various classes and types of CRISPR / Cas loci. [Figure 2] This diagram illustrates the structure of a natural class II, type II crRNA / tracrRNA pair, compared to hybrid sgRNAs formed by the ligation of crRNA and tracrRNA, as shown in Cas9 and other models. [Figure 3] This diagram illustrates the two paths observed in Tn7 and Tn7-like elements. [Figure 4] This figure illustrates the genomic context of Tn7CAST, type V of the MG64 family. The upper part of Figure 4A shows that the MG64-1CAST system consists of a CRISPR array (CRISPR repeats), a type V nuclease, and three predicted transposase protein sequences. TracrRNA was predicted in the intergenetic region between the CAST effector and the CRISPR array. The lower part of Figure 4A shows numerous sequence alignments of the catalytic domain of transposase TnsB. Catalytic residues are indicated by boxes. Figure 4B shows that two transposon ends are predicted for the MG64-1CAST system. [Figure 5] This figure depicts the predicted structure of the corresponding sgRNA of the CAST system described herein. Figure 5A (left) shows the predicted MG64-1 tracrRNA and crRNA double-stranded complex in a repeat-anti-repeat stem. The loop was truncated and a tetraloop of GAAA was added to the stem-loop structure to generate the designed sgRNA shown in Figure 5B (right). [Figure 6] This figure illustrates the results of a rearrangement reaction targeting a plasmid library consisting of NNNNNNNN at the 5' end of the target spacer sequence. Reaction #1 indicates the presence of the target library, #2 indicates the presence of donor fragments in both rearrangement reactions, and #3-#5 show sg-specific PCR bands corresponding to the appropriate rearrangement reactions. [Figure 7] This diagram illustrates the results of Sanger sequencing. Figure 7A shows the Sanger sequencing of the donor-target junction on the left end (LE) of the transposon when the LE is close to the PAM rearrangement reaction. The expected sequence, along with the predicted rearrangement event 61 bp away from PAM, is at the top of the panel. The chromatogram at the top is the result of sequencing originating from within the donor fragment. A clear signal is seen at the right end up to the donor / target junction (dotted line). This indicates mixing of rearrangement products. The chromatogram at the bottom of the panel is the sequencing from the target to the donor / target junction. The signal on the left is a clear signal up to the junction point. Figure 7B shows the Sanger sequencing of the donor-target junction on the right end (RE) of the transposon when the LE is close to the PAM product. The expected sequence, along with the predicted rearrangement event 61 bp away from PAM, is at the top of the panel. The chromatogram at the top is the result of sequencing originating from within the donor fragment. A clear signal is observed above the left edge up to the donor / target junction (dotted line). Figure 7C is a magnified view of the PAM library. Figure 7D is a SeqLogo analysis on NGS when the LE is close to the PAM event, showing a very strong preference for NGTNs in the PAM motif. [Figure 8] This diagram illustrates the phylogenetic gene tree of Cas12k effector sequences. This tree was inferred from multiple sequence alignments of 64 Cas12k sequences recovered in this study (orange and black branches) and 229 reference Cas12k sequences from publicly available databases (gray branches). Orange branches indicate Cas12k effectors whose association with CAST transposon components has been confirmed. [Figure 9]This figure shows the MG64 family CRISPR repeat alignment. The Cas12k CAST CRISPR repeat contains the conserved motif 5'-GNNGGNNTGAAAG-3'. In MG64-1, the short repeat-anti-repeat (RAR) within the CRISPR repeat motif aligns with the tracrRNA. The MG64RAR motif appears to define the start and end of the tracrRNA (5' end: RAR1(TTTC); 3' end: RAR2(CCNNC)). [Figure 10A] This figure shows the secondary structure predicted from the folding of CRISPR repeats + tracrRNA in relation to the MG64 system. [Figure 10B] This figure shows the secondary structure predicted from the folding of CRISPR repeats + tracrRNA in relation to the MG64 system. [Figure 11A] This diagram illustrates the MG64-3 CRISPR locus. TracrRNA is encoded upstream of the CRISPR array, and the transposon end is encoded downstream (inner black box). Sequences corresponding to partial 3' CRISPR repeats and partial spacers are encoded within the transposon (outer box). Self-matching spacers are encoded outside the transposon end. [Figure 11B] This diagram illustrates the tracrRNA sequence alignments for various CASTs provided herein. The tracrRNA sequence alignments indicate conserved regions. In particular, the sequence "TGCTTTC" (upper frame) at positions 92-98 is suggested to be important for the tertiary structure of sgRNA and for discontinuous repeat-anti-repeat pairing with crRNA. Furthermore, the hairpin "CYCC(n6)GGRG" (lower frame) at positions 265-278 is suggested to be functionally important and to allow for the positioning of downstream sequences for crRNA pairing. [Figure 12A] This shows the predicted structure of MG64-1 sgRNA. [Figure 12B] This shows the predicted structure of MG64-3 sgRNA. [Figure 12C] This shows the predicted structure of MG64-5 sgRNA. [Figure 13] This figure depicts PCR data demonstrating the activity of MG64-1 against sgRNA v2-1. The effector protein and its TnsB, TnsC, and TniQ proteins were expressed in an in vitro transcription / translation system using the protocol described for in vitro targeted integrase activity. Post-translation, target DNA, cargo DNA, and sgRNA were added to the reaction buffer. Integration was assayed by PCR across the target / donor junction. Figure 13A depicts a diagram illustrating the possible orientations of the integrated donor DNA. PCR reactions 3, 4, 5, and 6 represent the respective integration ligation products depending on the orientation of the donor integration at the target site. Figure 13B depicts a gel image of PCR 4 (detecting the RE junction to the donor) of transposition, showing: Lane 1) apo (no sgRNA), Lane 2) with sgRNA1, and Lane 3) with sgRNA v2-1. Figure 13C shows a gel image of PCR5 for translocation (detecting the LE junction to the donor), and shows the following: Lane 1) apo (no sgRNA), Lane 2) with sgRNA1, and Lane 3) with sgRNA v2-1. [Figure 14] This figure plots PCR reactions 5 (LE proximal to PAM, upper half of the plot) and 4 (RE distal to PAM, lower half of the plot) for MG64-1 on the sequence and distance from PAM. Analysis of the integration window showed that 95% of integrations occurring at the spacer PAM site occurred within a 10 bp window, 58–68 nucleotides away from PAM. The difference in integration distance between distal and proximal frequencies reflected overlaps of 3–5 base pairs at the integration site, resulting from shifts in transposase nuclease activity during integration. [Figure 15]This is a figure depicting the results of a colony PCR screen for translocation efficiency. After incubation, 18 colony forming units (CFUs) were seen on the plate, 8 on plate A (lanes labeled as A without IPTG), and 10 on plate B (lanes labeled as B with 100 μM IPTG at the time of recovery). All 18 were analyzed by colony PCR, which resulted in product bands showing an excellent translocation reaction (arrow). [Figure 16] This is a figure depicting the sequencing results of the selected colony PCR products, confirming that crossing the junction between LE and PAM at the engineered target site in the lacZ gene represents a translocation event. While the target and PAM are shown in gray, the minimal LE sequence is shown in blue at the top of the screen (minLE). A certain sequence variation is observed in the PCR products, which is expected if the insertion occurs at a variable distance upstream of the PAM. [Figure 17]Figure 17A shows the results of the manipulated single-guide test for 64-1 translocation activity. The black squares are lanes not related to this experiment. Figure 17A shows the gel image of PCR4 for translocation (detecting RE junctions to the donor). Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = sgRNA v1-1, Lane 4 = sgRNA v1-2, Lane 5 = sgRNA v1-3. Figure 17B shows the gel image of PCR5 for translocation (detecting LE junctions to the donor). Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = sgRNA v1-1, Lane 4 = sgRNA v1-2, Lane 5 = sgRNA v1-3. Figure 17C shows the gel image of PCR4 for translocation (detecting RE junctions to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = sgRNA v1-4, Lane 4 = sgRNA v1-6, Lane 5 = sgRNA v1-7, Lane 6 = sgRNA v1-8, Lane 7 = sgRNA v1-9. Figure 17D shows the gel image of PCR5 for translocation (detecting the LE junction to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = sgRNA v1-4, Lane 4 = sgRNA v1-6, Lane 5 = sgRNA v1-7, Lane 6 = sgRNA v1-8, Lane 7 = sgRNA v1-9. Figure 17E shows the gel image of PCR4 for translocation (detecting the RE junction to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = sgRNA v1-5, Lane 4 = Skip, Lane 5 = sgRNA v1-10. Figure 17F shows the gel image of PCR5 for translocation (detecting the LE junction to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = sgRNA v1-5, Lane 4 = Skip, Lane 5 = sgRNA v1-10. Figure 17G shows the gel image of PCR4 for translocation (detecting the RE junction to the donor).Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = sgRNA v1-17, Lane 4 = sgRNA v1-18, Lane 5 = Skip, Lane 6 = sgRNA v1-19, Lane 7 = Skip, Lane 8 = sgRNA v1-20. H in Figure 17 shows the gel image of PCR5 for translocation (detecting the LE junction to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = sgRNA v1-17, Lane 4 = sgRNA v1-18, Lane 5 = Skip, Lane 6 = sgRNA v1-19, Lane 7 = Skip, Lane 8 = sgRNA v1-20. [Figure 18]Figure 18-1 shows the results of the tests for manipulated LE and RE for translocation activity. The black squares are lanes not related to this experiment. Figure 18A shows the gel image of PCR4 for translocation (detecting the RE junction to the donor). Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = LE 86bp, Lane 4 = LE 105bp, Lane 5 = RE 196bp, Lane 6 = RE 242bp, Lane 7 = RE internal deletion 50, Lane 8 = RE internal deletion 81. Figure 18B shows the gel image of PCR5 for translocation (detecting the LE junction to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = LE 86bp, Lane 4 = LE 105bp, Lane 5 = RE 196bp, Lane 6 = RE 242bp, Lane 7 = RE internal deletion 50, Lane 8 = RE internal deletion 81. Figure 18C shows the gel image of PCR4 for translocation (detecting the RE junction to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = RE internal deletions 81 and 178bp, Lane 4 = skipped, Lane 5 = RE internal deletions 81 and 196bp, Lane 6 = skipped, Lane 7 = RE internal deletions 81 and 212bp, Lane 8 = skipped. Figure 18D shows the gel image of PCR5 for translocation (detecting the LE junction to the donor). Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = RE internal deletions 81 and 178 bp, Lane 4 = skipped, Lane 5 = RE internal deletions 81 and 196 bp, Lane 6 = skipped, Lane 7 = RE internal deletions 81 and 212 bp, Lane 8 = skipped. Figure 18 E shows the gel image of PCR4 for translocation (detecting the RE junction to the donor). Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = RE internal deletions 81 and 178 bp + LE 68 bp, Lane 4 = RE internal deletions 81 and 178 bp + LE 86 bp, Lane 5 = skipped, Lane 6 = RE internal deletions 81 and 178 bp + LE 105 bp, Lane 7 = skipped. Figure 18 F shows the gel image of PCR5 for translocation (detecting the LE junction to the donor).Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = RE internal deletion 81 and 178 bp + LE 68 bp, Lane 4 = RE internal deletion 81 and 178 bp + LE 86 bp, Lane 5 = skip, Lane 6 = RE internal deletion 81 and 178 bp + LE 105 bp, Lane 7 = skip. G in Figure 18 shows the gel image of transposition PCR6 (detecting the RE junction to the donor). Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = 0 bp overhang, Lane 4 = 1 bp overhang, Lane 5 = 2 bp overhang, Lane 6 = 3 bp overhang, Lane 7 = 5 bp overhang, Lane 8 = 10 bp overhang. [Figure 19]This figure shows the results of testing manipulated CAST components with NLS for translocation activity. The black squares are lanes not related to this experiment. Figure 19A shows the gel image of PCR4 for translocation (detecting RE junctions to the donor). Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = skipped, Lane 4 = skipped, Lane 5 = skipped, Lane 6 = NLS-TnsB, Lane 7 = skipped, Lane 8 = TnsB-NLS. Figure 19B shows the gel image of PCR5 for translocation (detecting LE junctions to the donor). Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = skipped, Lane 4 = skipped, Lane 5 = skipped, Lane 6 = NLS-TnsB, Lane 7 = skipped, Lane 8 = TnsB-NLS. Figure 19C shows the gel image of PCR4 for translocation (detecting RE junctions to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = Skip, Lane 4 = Skip, Lane 5 = Skip, Lane 6 = NLS-TniQ, Lane 7 = Skip, Lane 8 = TniQ-NLS. Figure 19D shows the gel image of PCR5 for translocation (detecting the LE junction to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = Skip, Lane 4 = Skip, Lane 5 = Skip, Lane 6 = NLS-TniQ, Lane 7 = Skip, Lane 8 = TniQ-NLS. Figure 19E shows the gel image of PCR4 for translocation (detecting the RE junction to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = Skip, Lane 4 = Skip, Lane 5 = NLS-Cas12k, Lane 6 = Cas12k-NLS, Lane 7 = NLS-TnsC, Lane 8 = TnsC-NLS. Figure 19F shows the gel image of PCR5 for translocation (detecting the LE junction to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = Skip, Lane 4 = Skip, Lane 5 = NLS-Cas12k, Lane 6 = Cas12k-NLS, Lane 7 = NLS-TnsC, Lane 8 = TnsC-NLS. Figure 19G shows the gel image of PCR4 for translocation (detecting the RE junction to the donor).Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = NLS-HA-TnsC, Lane 4 = NLS-TnsC-FLAG, Lane 5 = NLS-TnsC-HA, Lane 6 = NLS-TnsC-Myc, Lane 7 = NLS-FLAG-TnsC, Lane 8 = NLS-Myc-TnsC. H in Figure 19 shows the gel image of PCR5 for transposition (detecting the LE junction to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = NLS-HA-TnsC, Lane 4 = NLS-TnsC-FLAG, Lane 5 = NLS-TnsC-HA, Lane 6 = NLS-TnsC-Myc, Lane 7 = NLS-FLAG-TnsC, Lane 8 = NLS-Myc-TnsC. Figure 19, section I, shows the gel image of PCR4 for translocation (detecting the RE junction to the donor). Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = Cas2 × NLS apo (no sgRNA), Lane 4 = Cas2 × NLS holo (+sgRNA). Figure 19, section J, shows the gel image of PCR5 for translocation (detecting the LE junction to the donor). Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = Cas2 × NLS apo (no sgRNA), Lane 4 = Cas2 × NLS holo (+sgRNA). [Figure 20] This diagram shows manipulated CAST-NLS acting as a single set. Unless otherwise noted, all lanes contain Cas12k-NLS as well as NLS-TniQ, TnsB, TnsC, and sgRNA. Figure 20A shows the gel image of PCR4 for translocation (detecting RE junctions to the donor). Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = NLS-TnsB, Lane 4 = TnsB-NLS, Lane 5 = NLS-TnsB and NLS-TnsC, Lane 6 = TnsB-NLS and NLS-TnsC. Figure 20B shows the gel image of PCR5 for translocation (detecting LE junctions to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+ sgRNA), Lane 3 = NLS-TnsB, Lane 4 = TnsB-NLS, Lane 5 = NLS-TnsB and NLS-TnsC, Lane 6 = TnsB-NLS and NLS-TnsC. [Figure 21]This figure illustrates the results of testing Cas effector and TniQ protein fusions on translocation activity. Figure 21A shows the gel image of PCR4 for translocation (detecting the RE junction to the donor). Lane 1 = apo with Cas-TniQ fusion (no sgRNA), Lane 2 = holo with Cas-TniQ fusion (+sgRNA), Lane 3 = apo with TniQ-Cas fusion (no sgRNA), Lane 4 = holo with TniQ-Cas fusion (+sgRNA). Figure 21B shows the gel image of PCR5 for translocation (detecting the LE junction to the donor). Lane 1 = apo with Cas-TniQ fusion (no sgRNA), Lane 2 = holo with Cas-TniQ fusion (+sgRNA), Lane 3 = apo with TniQ-Cas fusion (no sgRNA), Lane 4 = holo with TniQ-Cas fusion (+sgRNA). Figure 21C shows the gel image of PCR4 for translocation (detecting the RE junction to the donor). Lane 1 = apo with TniQ-Cas fusion (no sgRNA), Lane 2 = holo with TniQ-Cas fusion (+sgRNA), Lane 3 = holo with Cas only, Lane 4 = apo with TniQ-48 linker-Cas fusion (no sgRNA), Lane 5 = holo with TniQ-48 linker-Cas fusion (+sgRNA), Lane 6 = apo with TniQ-68 linker-Cas fusion (no sgRNA), Lane 7 = holo with TniQ-68 linker-Cas fusion (+sgRNA), Lane 8 = holo with TniQ-72 linker-Cas fusion (+sgRNA). Figure 21D shows the gel image of PCR5 for translocation (detecting the LE junction to the donor). Lane 1 = Apo with TniQ-Cas fusion (no sgRNA), Lane 2 = Holo with TniQ-Cas fusion (+sgRNA), Lane 3 = Holo with Cas only, Lane 4 = Apo with TniQ-48 linker-Cas fusion (no sgRNA), Lane 5 = Holo with TniQ-48 linker-Cas fusion (+sgRNA), Lane 6 = Apo with TniQ-68 linker-Cas fusion (no sgRNA), Lane 7 = Holo with TniQ-68 linker-Cas fusion (+sgRNA), Lane 8 = Lane holo with TniQ-72 linker-Cas fusion (+sgRNA).Figure 21E shows the gel image of PCR4 for translocation (detecting the RE junction to the donor). Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = apo with NLS-TniQ-Cas-NLS fusion (no sgRNA), Lane 4 = holo with NLS-TniQ-Cas-NLS fusion (+sgRNA), Lane 5 = apo with NLS-TniQ-77 linker-Cas-NLS fusion (no sgRNA), Lane 6 = holo with NLS-TniQ-77 linker-Cas-NLS fusion (+sgRNA). Figure 21F shows the gel image of PCR5 for translocation (detecting the LE junction to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = Apo with NLS-TniQ-Cas-NLS fusion (no sgRNA), Lane 4 = Holo with NLS-TniQ-Cas-NLS fusion (+sgRNA), Lane 5 = Apo with NLS-TniQ-77 linker-Cas-NLS fusion (no sgRNA), Lane 6 = Holo with NLS-TniQ-77 linker-Cas-NLS fusion (+sgRNA). G in Figure 21 shows the gel image of PCR4 for transposition (detecting the RE junction to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = NLS-TniQ-Cas-NLS Apo (no sgRNA), Lane 4 = NLS-TniQ-Cas-NLS Holo (+sgRNA), Lane 5 = Cas-NLS-P2A-NLS-TniQ Apo (no sgRNA), Lane 6 = Cas-NLS-P2A-NLS-TniQ Holo (+sgRNA). H in Figure 21 shows the gel image of PCR5 for transposition (detecting the LE junction to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = NLS-TniQ-Cas-NLS Apo (no sgRNA), Lane 4 = NLS-TniQ-Cas-NLS Holo (+sgRNA), Lane 5 = Cas-NLS-P2A-NLS-TniQ Apo (no sgRNA), Lane 6 = Cas-NLS-P2A-NLS-TniQ Holo (+sgRNA). [Figure 22]This figure illustrates the results of TnsB and TnsC expression in human cells, followed by cell fractionation and in vitro translocation reactions. Figure 22A shows a gel image of PCR4 for translocation (detecting the RE junction to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = Untreated (no TnsB) holo with cytoplasm (+sgRNA), Lane 4 = Untreated holo with nucleoplasm (+sgRNA), Lane 5 = NLS-TnsB cell holo with cytoplasm (+sgRNA), Lane 6 = NLS-TnsB cell holo with nucleoplasm (+sgRNA), Lane 7 = TnsB-NLS cell holo with cytoplasm (+sgRNA), Lane 8 = TnsB-NLS cell holo with nucleoplasm (+sgRNA), Lane 9 = NLS-TniQ cell holo with cytoplasm (+sgRNA), Lane 10 = NLS-TniQ cell holo with nucleoplasm (+sgRNA). Figure 22B shows the gel image of PCR5 for translocation (detecting the LE junction to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = Holo with cytoplasm (+sgRNA) of untreated (no TnsB), Lane 4 = Holo with untreated nucleoplasm (+sgRNA), Lane 5 = Holo with cytoplasm (+sgRNA) of NLS-TnsB cells, Lane 6 = Holo with nucleoplasm (+sgRNA) of NLS-TnsB cells, Lane 7 = Holo with cytoplasm (+sgRNA) of TnsB-NLS cells, Lane 8 = Holo with nucleoplasm (+sgRNA) of TnsB-NLS cells, Lane 9 = Holo with cytoplasm (+sgRNA) of NLS-TniQ cells, Lane 10 = Holo with nucleoplasm (+sgRNA) of NLS-TniQ cells. C in Figure 22 shows the gel image of PCR4 for transposition (detecting the RE junction to the donor).Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = Holo without TnsC (+sgRNA), Lane 4 = Holo with untreated (no TnsC) nucleoplasm (+sgRNA), Lane 5 = Holo with untreated nucleoplasm (+sgRNA), Lane 6 = Holo with cytoplasm of NLS-HA-TnsC cells (+sgRNA), Lane 7 = Holo with nucleoplasm of NLS-HA-TnsC cells (+sgRNA), Lane 8 = Holo with cytoplasm of TnsC-NLS cells (+sgRNA), Lane 9 = Holo with nucleoplasm of TnsC-NLS cells (+sgRNA). D in Figure 22 shows the gel image of PCR5 for transposition (detecting the LE junction to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = Holo without TnsC (+sgRNA), Lane 4 = Holo with untreated (no TnsC) nucleoplasm (+sgRNA), Lane 5 = Holo with untreated nucleoplasm (+sgRNA), Lane 6 = Holo with cytoplasm of NLS-HA-TnsC cells (+sgRNA), Lane 7 = Holo with nucleoplasm of NLS-HA-TnsC cells (+sgRNA), Lane 8 = Holo with cytoplasm of TnsC-NLS cells (+sgRNA), Lane 9 = Holo with nucleoplasm of TnsC-NLS cells (+sgRNA). Figure 22E shows the gel image of PCR4 for transposition (detecting the RE junction to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = Apo (no sgRNA) NLS-TnsB-IRES-NLS-TnsC cytoplasm, Lane 4 = Holo (+sgRNA) NLS-TnsB-IRES-NLS-TnsC cytoplasm, Lane 5 = Apo (no sgRNA) NLS-TnsB-IRES-NLS-TnsC nucleoplasm, Lane 6 = Holo (+sgRNA) NLS-TnsB- Lane 7 = IRES-NLS-TnsC nucleoplasm, Lane 8 = Apo (no sgRNA) TnsB-NLS-IRES-NLS-TnsC cytoplasm, Lane 9 = Apo (no sgRNA) TnsB-NLS-IRES-NLS-TnsC nucleoplasm, Lane 10 = Holo (+sgRNA) TnsB-NLS-IRES-NLS-TnsC nucleoplasm. F in Figure 22 shows the gel image of PCR5 for transposition (detecting the LE junction to the donor).Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = Apo (no sgRNA) NLS-TnsB-IRES-NLS-TnsC cytoplasm, Lane 4 = Holo (+sgRNA) NLS-TnsB-IRES-NLS-TnsC cytoplasm, Lane 5 = Apo (no sgRNA) NLS-TnsB-IRES-NLS-TnsC nucleoplasm, Lane 6 = Holo (+sgRNA) NLS-TnsB- IRES-NLS-TnsC nucleoplasm, lane 7 = apo (no sgRNA) TnsB-NLS-IRES-NLS-TnsC cytoplasm, lane 8 = holo (+sgRNA) TnsB-NLS-IRES-NLS-TnsC cytoplasm, lane 9 = apo (no sgRNA) TnsB-NLS-IRES-NLS-TnsC nucleoplasm, lane 10 = holo (+sgRNA) TnsB-NLS-IRES-NLS-TnsC nucleoplasm. [Figure 23]This figure shows the results of the expression of Cas12k and TniQ-linked components in human cells, followed by an in vitro translocation test. Figure 23A shows the gel image of PCR5 for translocation (detecting the LE junction to the donor). Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = Cas-NLS holo (+sgRNA) cytoplasm, Lane 4 = Cas-NLS holo (+sgRNA) nucleoplasm, Lane 5 = Cas-NLS holo (+sgRNA) nucleoplasm + additional sgRNA, Lane 6 = Cas-NLS-P2A-NLS-TniQ holo (+sgRNA) cytoplasm, Lane 7 = Cas-NLS-P2A-NLS-TniQ holo (+sgRNA) nucleoplasm, Lane 8 = Cas-NLS-P2A-NLS-TniQ holo (+sgRNA) nucleoplasm + additional sgRNA. Figure 23B shows the gel image of PCR4 for translocation (detecting the RE junction site to the donor). Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = apo (no sgRNA) Cas-NLS-P2A-NLS-TniQ cytoplasm, Lane 4 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ cytoplasm, Lane 5 = apo (no sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm, Lane 6 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm, Lane 7 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ cytoplasm + additional holo-Cas-NLS, Lane 8 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm + NLS-TniQ. Figure 23C shows a gel image of PCR5 (detecting the LE junction to the donor) of the dislocation.Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = Apo (no sgRNA) Cas-NLS-P2A-NLS-TniQ cytoplasm, Lane 4 = Holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ cytoplasm, Lane 5 = Apo (no sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm, Lane 6 = Holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm, Lane 7 = Holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ cytoplasm + additional holo-Cas-NLS, Lane 8 = Holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm + NLS-TniQ. D in Figure 23 shows the gel image of PCR4 for transposition (detecting the RE junction to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = Apo (no sgRNA) NLS-TniQ-Cas-NLS cytoplasm, Lane 4 = Holo (+sgRNA) NLS-TniQ-Cas-NLS cytoplasm, Lane 5 = Apo (no sgRNA) NLS-TniQ-Cas-NLS nucleoplasm, Lane 6 = Holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm, Lane 7 = Holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm + additional HoloCas-NLS, Lane 8 = Holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm + NLS-TniQ. Figure 23E shows the gel image of PCR5 for transposition (detecting the LE junction to the donor). Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = Apo (no sgRNA) NLS-TniQ-Cas-NLS cytoplasm, Lane 4 = Holo (+sgRNA) NLS-TniQ-Cas-NLS cytoplasm, Lane 5 = Apo (no sgRNA) NLS-TniQ-Cas-NLS nucleoplasm, Lane 6 = Holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm, Lane 7 = Holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm + additional HoloCas-NLS, Lane 8 = Holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm + NLS-TniQ. F in Figure 23 shows the gel image of PCR4 for transposition (detecting the RE junction to the donor).Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = Apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ cytoplasm, Lane 4 = Holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ cytoplasm, Lane 5 = Apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm, Lane 6 = Apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional PURExpress, Lane 7 = Apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional Lane 8 = Apo(no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + NLS-TniQ, Lane 9 = Holo(+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm, Lane 10 = Holo(+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional PURExpress, Lane 11 = Holo(+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional Cas-NLS, Lane 12 = Holo(+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + NLS-TniQ. G in Figure 23 shows the gel image of PCR5 for transposition (detecting the LE junction to the donor).Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = Apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ cytoplasm, Lane 4 = Holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ cytoplasm, Lane 5 = Apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm, Lane 6 = Apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional PURExpress, Lane 7 = Apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional Cas-NLS, lane 8 = apo(no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + NLS-TniQ, lane 9 = holo(+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm, lane 10 = holo(+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional PURExpress, lane 11 = holo(+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional Cas-NLS, lane 12 = holo(+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + NLS-TniQ. [Figure 24] This figure illustrates the electrophoretic transfer assay (EMSA) results for 64-1 TnsB and its LE DNA sequence. EMSA results confirm binding and TnsB recognition. TnsB protein was expressed in an in vitro transcription / translation system, incubated with FAM-labeled DNA containing the LE sequence, and separated on a natural 5% TBE gel. Binding is observed as an upward shift of the labeled band. Due to the presence of multiple TnsB binding sites, multiple shifts are observed in the EMSA. Lane 1: FAM-labeled DNA only. Lane 2: FAM DNA and in vitro transcription / translation system (without TnsB protein). Lane 3: FAM DNA and TnsB.

[0013] A brief description of the sequence listing The sequence listings submitted with this specification provide exemplary polynucleotide and polypeptide sequences for use in the methods, compositions, and systems relating to this disclosure. The following is an exemplary description of some of these sequences.

[0014] MG64

[0015] Sequence IDs 1, 12, 16, 20-30, 64, and 80-85 show the full-length peptide sequences of the MG64Cas effector.

[0016] Sequence IDs 2-4, 13-15, 17-19, and 65-67 show peptide sequences of MG64 transposition proteins that may contain a recombinase complex associated with the MG64Cas effector.

[0017] Sequence IDs 5-6, 32-33, 94-95, and 104-105 show the nucleotide sequences of MG64tracrRNA derived from the same locus as the MG64Cas effector.

[0018] Sequence IDs 7 and 34-35 show the nucleotide sequences of the MG64-targeted CRISPR repeat.

[0019] Sequence IDs 106-108 show the nucleotide sequences of MG64crRNA.

[0020] Sequence IDs 8, 10, 39-44, 77, 79, and 93 show the nucleotide sequences of the right-hand transposase recognition sequences associated with the MG64 system.

[0021] Sequence IDs 9, 11, 36-38, 76, and 78 show the nucleotide sequences of the left-hand transposase recognition sequences associated with the MG64 system.

[0022] Sequence ID 31 shows the PAM sequence associated with the MG64Cas effect pedal described herein.

[0023] Sequence IDs 45–63, 68–75, and 96–103 show the nucleotide sequences of single guide RNAs engineered to function with the MG64Cas effector.

[0024] Other arrays Sequence IDs 86-87 show the peptide sequences of the nuclear localization signal.

[0025] Sequence IDs 88-89 show the linker peptide sequences.

[0026] Sequence IDs 90-92 show the peptide sequences of the epitope tag. [Modes for carrying out the invention]

[0027] While embodiments of the present invention are shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided only as examples. It will be understood by those skilled in the art that numerous modifications, variations, and substitutions can be made without departing from the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein may be utilized.

[0028] The implementation of some of the methods disclosed herein utilizes techniques from immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, unless otherwise specified. For example, see Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (FMAusubel, et al. eds.); the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (MJ MacPherson, BD Hames and GRTaylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (RIFreshney, ed. (2010)) (which are fully incorporated herein by reference).

[0029] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context otherwise expressly indicates. Furthermore, to the extent that the terms "including," "includes," "having," "has," "with," or their variations thereof are used in any of the detailed description and / or claims, such terms are intended to be inclusive in a manner similar to that of the term "comprising."

[0030] The terms "about" or "approximately" mean that a particular value is within an acceptable margin of error as determined by those skilled in the art, and this depends in part on how that value is measured or determined, i.e., on the limitations of the measurement system. For example, "about" may mean a standard deviation of 1 or more for the practice in the art. Alternatively, "about" may mean a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of any given value.

[0031] As used herein, “cell” generally refers to a living cell. A cell can be the basic structural, functional, and / or biological unit of a living organism. A cell may originate from any organism that has one or more cells. Some non-limiting examples include prokaryotic cells, eukaryotic cells, bacterial cells, archaeal cells, cells of unicellular eukaryotes, protist cells, plant cells (e.g., crops, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkins, hay, potatoes, cotton, cannabis, tobacco, flowering plants, conifers, gymnosperms, ferns, clubmosses, hornworts, liverworts, mosses), algal cells (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens). Examples include cells from organisms such as C. agardh, seaweed (e.g., kelp), fungal cells (e.g., yeast cells, mushroom cells), animal cells, invertebrate cells (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), vertebrate cells (e.g., fish, amphibians, reptiles, birds, mammals), and mammalian cells (e.g., pigs, cows, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.). Cells are those that do not originate from naturally occurring organisms (e.g., synthetically produced cells, often called artificial cells).

[0032] The term "nucleotide," as used herein, generally refers to a combination of base-sugar-phosphate. Nucleotides may include synthetic nucleotides. Nucleotides may include synthetic nucleotide analogs. Nucleotides can be monomeric units of nucleic acid sequences (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide may include ribonucleoside triphosphates such as adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP), and deoxyribonucleoside triphosphates, e.g., dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives may include, for example, [αS]dATP, 7-deaza-dGTP, and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. As used herein, the term nucleotide may also refer to dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Exemplary examples of dideoxyribonucleoside triphosphates include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides may be unlabeled or labeled in a detectable manner, for example, by using a moiety containing an optically detectable moiety (e.g., a fluorophore). Labeling may also be performed using quantum dots. Examples of detectable labels include radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzymatic labels. Examples of fluorescent labels for nucleotides include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxylrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxylrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4'-dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP; FluoroLink DeoxyNucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP available from Perkin Elmer (Foster City, Calif); Boehringer Fluorescein-15-dATP, fluorescein-12-dUTP, tetramethylrhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2'-dATP, available from Mannheim (Indian Apolis, Ind.); as well as Molecular Available chromosome-labeled nucleotides from Probes (Eugene, Oreg.) include BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP. Nucleotides can be further labeled or marked by chemical modifications.A chemically modified single nucleotide can be a biotin-dNTP. Some non-limiting examples of biotinylated dNTPs include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).

[0033] The terms “polynucleotide,” “oligonucleotide,” and “nucleic acid” are generally used interchangeably to mean polymeric forms of nucleotides of any length, whether single-stranded, double-stranded, or multi-stranded, deoxyribonucleotides or ribonucleotides, or their analogues. Polynucleotides may be exogenous or endogenous to cells. Polynucleotides may exist in non-cellular environments. Polynucleotides may be genes or fragments thereof. Polynucleotides may be DNA. Polynucleotides may be RNA. Polynucleotides may have any three-dimensional structure and may perform any function. Polynucleotides may contain one or more analogues (e.g., altered backbone, sugar, or nucleic acid base). Modifications to the nucleotide structure, if present, may be given before or after polymer assembly. Some non-exclusive examples of analogs include 5-bromouracil, peptide nucleic acids, xeno nucleic acids, morpholino, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., sugar-bound rhodamine or fluorescein), thiol-containing nucleotides, biotin-bound nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, keosin, and waiosin. Non-limiting examples of polynucleotides include coding or non-coding regions of genes or gene fragments, loci (multiple loci) defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), small interfering RNA (siRNA), small hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers.The sequence of nucleotides may be interrupted by non-nucleotide components.

[0034] The terms "transfection" or "transfected" generally refer to the introduction of nucleic acids into cells by non-viral or virus-based methods. Nucleic acid molecules may be gene sequences encoding complete proteins or their functional portions. See, for example, Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, 18.1–18.88.

[0035] The terms “peptide,” “polypeptide,” and “protein” are used interchangeably herein to generally refer to polymers of at least two amino acid residues linked by peptide bonds. This term does not imply a specific length of polymer, nor is it intended to suggest or distinguish whether a peptide is produced using recombinant techniques, chemical or enzymatic synthesis, or naturally occurring. This term applies to naturally occurring amino acid polymers as well as amino acid polymers containing at least one modified amino acid. In some cases, the polymer may be interrupted by non-amino acids. This term includes amino acid chains of any length, including full-length proteins, as well as proteins with or without secondary and / or tertiary structures (e.g., domains). This term further encompasses amino acid polymers modified by any other operations, such as disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, oxidation, and conjugation with labeling components. The terms “amino acid” and “amino acid” generally refer, when used herein, to naturally occurring and unnatural amino acids, including but not limited to modified amino acids and amino acid analogs. Modified amino acids may include both natural and unnatural amino acids that have been chemically modified to include a group or chemical moiety that does not naturally exist on the amino acid surface. Amino acid analogs may also refer to amino acid derivatives. The term "amino acid" includes both D-amino acids and L-amino acids.

[0036] As used herein, “unnatural” can generally refer to a sequence of nucleic acid or polypeptide not found in natural nucleic acids or proteins. “Unnatural” may refer to an affinity tag. “Unnatural” may refer to a fusion. “Unnatural” may refer to a sequence of naturally occurring nucleic acid or polypeptide including mutations, insertions, and / or deletions. A nonnatural sequence may exhibit and / or encode activity (e.g., enzymatic activity, methyltransferase activity, acetyltransferase activity, kinase activity, ubiquitination activity, etc.) that may also be exhibited by the nucleic acid and / or polypeptide sequence to which the nonnatural sequence is fused. A nonnatural nucleic acid or polypeptide sequence may be ligated to a naturally occurring nucleic acid or polypeptide sequence (or a variant thereof) by genetic engineering to produce a chimeric nucleic acid and / or polypeptide sequence encoding a chimeric nucleic acid and / or polypeptide.

[0037] The term “promoter” generally refers, as used herein, to a regulatory DNA region that controls the transcription or expression of a gene and may be located adjacent to or overlapping with a nucleotide or region of nucleotides from which RNA transcription is initiated. Promoters may often contain specific DNA sequences that bind to protein factors called transcription factors, facilitating the binding of RNA polymerase to the DNA that leads to gene transcription. A “basic promoter,” also called a “core promoter,” may generally refer to a promoter that contains all the fundamental elements necessary to promote the transcription and expression of a functionally linked polynucleotide. Eukaryotic basic promoters typically, though not always, contain a TATA-box and / or CAAT-box.

[0038] The term “expression,” as used herein, generally refers to the process by which a nucleic acid sequence or polynucleotide is transcribed from a DNA template (into mRNA or other RNA transcripts, etc.), and / or the process by which the transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. The transcript and the encoded polypeptide are sometimes collectively referred to as “gene products.” When the polynucleotide originates from genomic DNA, expression may involve the splicing of mRNA in eukaryotic cells.

[0039] As used herein, “operably linked,” “operably linked,” “operably linked,” or their grammatical equivalents generally refer to the juxtaposition of genetic elements, such as promoters, enhancers, and polyadenylation sequences, which are related in a way that enables them to function in a desired manner. For example, a regulatory element, which may include a promoter and / or enhancer sequence, is operably linked to a coding region if the regulatory element helps initiate transcription of the coding sequence. Intervening residues may exist between the regulatory element and the coding region, as long as this functional relationship is maintained.

[0040] When used herein, "vector" generally refers to a macromolecule or association of macromolecules that contains or associates with polynucleotides and may be used to mediate the delivery of polynucleotides to cells. Examples of vectors include plasmids, viral vectors, liposomes, and other gene delivery vehicles. Vectors generally include genetic elements, such as regulatory elements, that are operably ligated to a gene to promote gene expression at a target.

[0041] As used herein, “expression cassette” and “nucleic acid cassette” are generally used interchangeably to refer to a combination of nucleic acid sequences or elements that are expressed together or operably linked for expression. In some cases, an expression cassette refers to regulatory elements and a combination of genes or genes to which they are operably linked for expression.

[0042] A “functional fragment” of a DNA or protein sequence generally refers to a fragment that possesses biological activity (either functional or structural) substantially similar to the biological activity of the full-length DNA or protein sequence. The biological activity of a DNA sequence may be its ability to influence expression in ways known to be attributable to the full-length sequence.

[0043] As used herein, the term "engineered" generally refers to an object that has been modified by human intervention. In a non-limiting example, a nucleic acid may be modified by altering its sequence to one that does not exist in nature. A nucleic acid may also be modified by ligating it with a nucleic acid that does not associate in nature, so that the ligated product has a function not present in the original nucleic acid. Engineered nucleic acids may be synthesized in vitro with sequences that do not exist in nature. A protein may be modified by altering its amino acid sequence to one that does not exist in nature. Engineered proteins may acquire new functions or properties. An "engineered" system contains at least one engineered component.

[0044] As used herein, “synthetic” and “artificial” are interchangeable to refer to proteins or domains that have low sequence identity (e.g., less than 50%, less than 25%, less than 10%, less than 5%, less than 1%) relative to naturally occurring human proteins. For example, the VPR and VP64 domains are synthetic transactivation domains.

[0045] The terms "tracrRNA" or "tracr sequence" generally refer, when used herein, to exemplary wild-type tracrRNA sequences (e.g., tracrRNA derived from S. pyogenes, S. aureus, etc., or sequence numbers: * _ * A tracrRNA can refer to a nucleic acid having at least approximately 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% sequence identity and / or similarity to a wild-type exemplary tracrRNA sequence (e.g., tracrRNA derived from S. pyogenes, S. aureus, etc.). A tracrRNA can also refer to a modified form of tracrRNA that may include nucleotide changes such as deletions, insertions, substitutions, mutations, mutations, or chimeras. A tracrRNA may refer to a nucleic acid that is at least approximately 60% identical to a wild-type exemplary tracrRNA sequence (e.g., tracrRNA derived from S. pyogenes, S. aureus, etc.) over a stretch of at least six consecutive nucleotides. For example, a tracrRNA sequence may be at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 95%, at least approximately 98%, at least approximately 99%, or 100% identical to a wild-type exemplary tracrRNA sequence (e.g., tracrRNA derived from S. pyogenes, S. aureus, etc.) over a stretch of at least six consecutive nucleotides. Type II tracrRNA sequences can be predicted on a genomic sequence by identifying regions that are complementary to parts of the repeat sequences of an adjacent CRISPR array.

[0046] As used herein, “guide nucleic acid” can generally refer to a nucleic acid that can hybridize to another nucleic acid. The guide nucleic acid may be RNA. The guide nucleic acid may be DNA. The guide nucleic acid may be programmed to bind site-specifically to a nucleic acid sequence. The target nucleic acid, i.e., the target nucleic acid, may contain nucleotides. The guide nucleic acid may contain nucleotides. Part of the target nucleic acid may be complementary to part of the guide nucleic acid. A double-stranded target polynucleotide chain that is complementary to the guide nucleic acid and hybridizes with the guide nucleic acid may be called a complementary chain. A double-stranded target polynucleotide chain that is complementary to the complementary chain and therefore may not be complementary to the guide nucleic acid may be called a non-complementary chain. The guide nucleic acid may contain a polynucleotide chain and may be called a “single guide nucleic acid”. The guide nucleic acid may contain two polynucleotide chains and may be called a “double guide nucleic acid”. Unless otherwise specified, the term “guide nucleic acid” may be comprehensive and refer to both single guide nucleic acids and double guide nucleic acids. The guide nucleic acid may contain a segment that can be called a “nucleic acid targeting segment” or “nucleic acid targeting sequence”. The nucleic acid targeting segment may include subsegments that are sometimes called "protein-binding segments," "protein-binding sequences," or "Cas protein-binding segments."

[0047] In the context of two or more nucleic acid or polypeptide sequences, the terms “sequence identity” or “percent identity” generally refer to two (e.g., in pairwise alignment) or more (e.g., in multiple sequence alignment) sequences that are identical or have a specific proportion of identical amino acid residues or nucleotides when compared and aligned for the greatest correspondence across a local or global comparison window, as measured using a sequence comparison algorithm. Suitable sequence comparison algorithms for polypeptide sequences include, for example, BLASTP using parameters of the BLOSUM62 scoring matrix with a word length (W) of 3, an expected value (E) of 10, an existence of 11, and an extension of 1, and conditional configuration score matrix adjustments for polypeptide sequences longer than 30 residues; BLASTP using the PAM30 scoring matrix with a word length (W) of 2, an expected value (E) of 1,000,000, and for sequences less than 30 residues, setting the gap cost to open a gap to 9 and the gap cost to extend a gap to 1 (these are the default parameters for BLASTP in the BLAST suite available at https: / / blast.ncbi.nlm.nih.gov); CLUSTALW using parameters of the Smith-Waterman homology search algorithm with parameters of 2 matches, -1 mismatches, and -1 gaps; MUSCLE with default parameters; MAFFT with parameters of 2 retrees and 1000 maxiterations; Novafold with default parameters; and HMMER with default parameters. hmmalign is one example.

[0048] This disclosure includes variants of any of the enzymes described herein that have one or more conserved amino acid substitutions. Such conserved substitutions can be made in the amino acid sequence of a polypeptide without disrupting the three-dimensional structure or function of the polypeptide. Conservative substitutions can be achieved by substituting amino acids that have similar hydrophobicity, polarity, and R-chain length. Furthermore or alternatively, conserved substitutions can be identified by comparing aligned sequences of homologous proteins from different species, thereby pinpointing amino acid residues that have mutated between species (e.g., non-conserved residues) without altering the basic function of the encoded protein. Such conservatively substituted mutants may include mutants having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of the systems described herein (e.g., the MG64 system described herein). In some embodiments, such conservatively substituted mutants are functional mutants. Such functional mutants may include sequences with substitutions such that the activity of key active site residues of the endonuclease is not disrupted. In some embodiments, any functional variant of any of the systems described herein lacks at least one substitution of the conserved or functional residues referred to in Figures 4 and 5. In some embodiments, any functional variant of any of the systems described herein lacks all substitutions of the conserved or functional residues referred to in Figures 4 and 5.

[0049] Tables of conserved substitutions that provide functionally similar amino acids are available from various sources (see, for example, Creighton, Proteins: Structures and Molecular Properties (WH Freeman & Co.; 2nd Edition (December 1993))). The following eight groups each contain amino acids that are conserved substitutions with each other. 1) Alanine (A), Glycine (G), 2) Aspartic acid (D), glutamic acid (E), 3) Asparagine (N), glutamine (Q), 4) Arginine (R), Lysine (K), 5) Isoleucine (I), leucine (L), methionine (M), valine (V), 6) Phenylalanine (F), tyrosine (Y), tryptophan (W), 7) Serine (S), threonine (T), and 8) Cysteine ​​(C), Methionine (M).

[0050] As used herein, the term “RuvC_III domain” generally refers to the third discontinuous segment of the RuvC endonuclease domain (the RuvC nuclease domain consists of three discontinuous segments: RuvC_I, RuvC_II, and RuvC_III). The RuvC domain or its segments can generally be identified by alignment to a known domain sequence, structural alignment to a protein with an annotated domain, or comparison with a hidden Markov model (HMM) constructed based on a known domain sequence (e.g., Pfam HMM PF18541 for RuvC_III).

[0051] As used herein, the term “HNH domain” generally refers to an endonuclease domain having characteristic histidine and asparagine residues. HNH domains can generally be identified by alignment to a known domain sequence, structural alignment to a protein with an annotated domain, or comparison with a hidden Markov model (HMM) constructed based on a known domain sequence (e.g., Pfam HMM PF01844 for the HNH domain).

[0052] As used herein, the term “recombinase” generally refers to a site-specific enzyme that mediates DNA recombination between recombinase-recognized sequences, resulting in the excision, integration, inversion, or exchange (e.g., translocation) of DNA fragments between recombinase-recognized sequences.

[0053] As used herein, the terms “recombinate” or “recombinant” in the context of nucleic acid modification (e.g., genome modification) generally refer to the process by which two or more nucleic acid molecules, or two or more regions of a single nucleic acid molecule, are modified by the action of a recombinase protein. Recombination can, in particular, result in, for example, the insertion, inversion, excision, or translocation of nucleic acid sequences within or between one or more nucleic acid molecules.

[0054] As used herein, the term “transposon” generally refers to a mobile element that moves in and out of the genome accompanied by “cargo DNA.” In some cases, these transposons may differ in the type of nucleic acid they transpose, the type of repeats at the end of the transposon, the type of cargo they carry, or the mode of transposition (i.e., self-repair or host repair). As used herein, “transposase” generally refers to an enzyme that binds to the end of a transposon and catalyzes its transposition to another part of the genome. In some cases, this transposition may be by a cut-and-paste mechanism or by a replication transposition mechanism.

[0055] As used herein, the terms “Tn7” or “Tn7-like transposase” generally refer to a family of transposases comprising three main components: heteromeric transposases (TnsA and / or TnsB) and regulatory proteins (TnsC). In addition to the TnsABC transposition proteins, the Tn7 element can encode dedicated target site-selective proteins, TnsD and TnsE. In addition to TnsABC, TnsD, a sequence-specific DNA-binding protein, directs transposition to a conserved site called the “Tn7 attachment site” (attTn7). TnsD is a member of a larger protein family that also includes TniQ. TniQ has been shown to target transposition to plasmid degrading sites.

[0056] In some cases, the CAST system described herein may include one or more Tn7 or Tn7-like transposases. In certain exemplary embodiments, the Tn7 or Tn7-like transposase includes a multimeric protein complex. In certain exemplary embodiments, the multimeric protein complex includes TnsA, TnsB, TnsC, or TniQ. In these combinations, the transposases (TnsA, TnsB, TnsC, TniQ) may form complexes or fusion proteins with each other.

[0057] As used herein, the term "Cas12k" (or alternatively, "Class II, Type VK") generally refers to a subtype of the Type V CRISPR system that has been found to be defective in nuclease activity (for example, they may contain at least one defective RuvC domain lacking at least one catalytic residue important for DNA cleavage). Effectors of such subtypes have generally been associated with the CAST system.

[0058] overview

[0059] The discovery of novel Cas enzymes with unique functionalities and structures could further disrupt deoxyribonucleic acid (DNA) editing technologies, offering the potential to improve speed, specificity, functionality, and ease of use. While the widespread adoption of CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) systems in microorganisms and indeed a vast array of microbial species is anticipated, the literature contains relatively few functionally characterized CRISPR / Cas enzymes. This is partly due to the sheer number of microbial species, which are not easily cultured in the laboratory. Metagenomic sequencing from natural environmental niches representing many microbial species could dramatically increase the number of known new CRISPR / Cas systems and accelerate the discovery of new oligonucleotide editing functions. A recent example demonstrating the usefulness of such an approach is the discovery of the CasX / CasY CRISPR system in 2016 from metagenomic analysis of natural microbial communities.

[0060] The CRISPR / Cas system is an RNA-directed nuclease complex described as functioning as an adaptive immune system in microorganisms. In its natural context, the CRISPR / Cas system occurs at the CRISPR (clustered regularly interspaced short palindromic repeats) operon or locus, which generally consists of two parts: (i) an array of short repetitive sequences (30-40 bp) also delimited by short spacer sequences that encode an RNA-based targeting element, and (ii) an ORF encoding Cas, which encodes a nuclease polypeptide directed by the RNA-based targeting element together with an accessory protein / enzyme. Efficient nuclease targeting of a particular target nucleic acid sequence generally requires both (i) complementary hybridization between the first 6-8 nucleic acids of the target (target seed) and the crRNA guide, and (ii) the presence of a protospacer-adjacent motif (PAM) sequence in a defined vicinity of the target seed (PAMs are typically sequences not commonly represented in the host genome). Depending on the precise function and organization of the system, CRISPR-Cas systems are generally classified into two classes, five types, and sixteen subtypes based on shared functional characteristics and evolutionary similarities (see Figure 1).

[0061] Class I CRISPR-Cas systems feature large-scale multi-subunit effector complexes and include types I, III, and IV.

[0062] Type I CRISPR-Cas systems are considered to have moderate complexity in terms of components. In type I CRISPR-Cas systems, the array of RNA targeting elements is transcribed as a long precursor crRNA (precrRNA), which, upon processing with repeat elements, releases a short, mature crRNA. The crRNA then directs the nuclease complex to the nucleic acid target when followed by a suitable short consensus sequence called a protospacer-adjacent motif (PAM). This processing occurs via the endoribonuclease subunit (Cas6) of a large endonuclease complex called a cascade, which also contains the nuclease (Cas3) protein component of the crRNA-directing nuclease complex. CasI nucleases primarily function as DNA nucleases.

[0063] Type III CRISPR systems are sometimes characterized by the presence of a central nuclease known as Cas10, along with a repeat-associated mysterious protein (RAMP) containing a Csm or Cmr protein subunit. Similar to type I systems, mature crRNA is processed from precrRNA using a Cas6-like enzyme. Unlike type I and type II systems, type III systems appear to target and cleave DNA-RNA double helixes (such as the DNA strand used as a template for RNA polymerase).

[0064] Type IV CRISPR-Cas systems have an effector complex consisting of a highly reduced large subunit nuclease (csf1), two genes for the RAMP protein in the Cas5 (csf3) and Cas7 (csf2) groups, and, in some cases, a gene for a predicted small subunit. Such systems are generally found on endogenous plasmids.

[0065] Type II CRISPR-Cas systems generally have a single polypeptide multi-domain nuclease effector and include types II, V, and VI.

[0066] Type II CRISPR-Cas systems are considered the simplest in terms of components. In type II CRISPR-Cas systems, processing of a CRISPR array with mature crRNA does not require the presence of a special endonuclease subunit, but rather a small transcoded crRNA (tracrRNA) with a region complementary to the array repeat sequence. The tracrRNA interacts with both the corresponding effector nuclease (e.g., Cas9) and the repeat sequence to form a precursor dsRNA structure, which is then cleaved by endogenous RNAseIII to produce a mature effector enzyme loaded with both tracrRNA and crRNA. CasII nucleases are known as DNA nucleases. Type II effectors generally exhibit a structure consisting of a RuvC-like endonuclease domain employing an RNase H fold and an inserted, unrelated HNH nuclease domain inserted within the fold of the RuvC-like nuclease domain. The RuvC-like domain is responsible for cleaving the target DNA strand (e.g., the complementary DNA strand of crRNA), while the HNH domain is responsible for cleaving the displaced DNA strand.

[0067] Type V CRISPR-Cas systems feature nuclease effector structures (e.g., Cas12) similar to those of type II effectors, including a RuvC-like domain. Like type II, most (but not all) type V CRISPR systems use tracrRNA to process precrRNA into mature crRNA. However, unlike type II systems which require RNAseIII to cleave precrRNA into multiple crRNAs, type V systems can use the effector nuclease itself to cleave precrRNA. Like type II CRISPR-Cas systems, type V CRISPR-Cas systems are again known as DNA nucleases. Unlike type II CRISPR-Cas systems, some type V enzymes (e.g., Cas12a) appear to possess robust single-strand nonspecific deoxyribonuclease activity activated by initial crRNA-directed cleavage of a double-stranded target sequence.

[0068] Type VI CRIPSR-Cas systems possess RNA-guided RNA endonucleases. Instead of a RuvC-like domain, single polypeptide effectors of type VI systems (e.g., Cas13) contain two HEPN ribonuclease domains. Unlike both type II and type V systems, type VI systems also appear not to require tracrRNA to process precrRNA to crRNA. However, similar to type V systems, some type VI systems (e.g., C2C2) appear to possess robust single-strand nonspecific nuclease (ribonuclease) activity activated by initial crRNA-directed cleavage of target RNA.

[0069] Due to their simpler architecture, Class II CRISPR-Cas are the most widely adopted for operation and development as designer nucleases / genome editing applications.

[0070] One of the earliest adaptations of such a system for in vitro use can be seen in Jinek et al. (Science. 2012 Aug 17;337(6096):816-21, fully invoked herein by reference). Jinek's study first used (i) recombinantly expressed, purified full-length Cas9 (e.g., class II, type II Cas enzyme) isolated from S. pyogenes SF370, (ii) purified mature crRNA of approximately 42 nt having an approximately 20 nt 5' sequence complementary to the target DNA sequence to be cleaved, followed by a 3' tracr binding sequence (the entire crRNA was transcribed in vitro from a synthetic DNA template carrying a T7 promoter sequence), (iii) purified tracrRNA transcribed in vitro from a synthetic DNA template carrying a T7 promoter sequence, and (iv) Mg 2+He described a system that included (ii) and (iii). Jinek later described an improved, engineered system in which the crRNA of (ii) is bound to the 5' end of (iii) by a linker (e.g., GAAA) to form a single fusion synthetic guide RNA (sgRNA) that can itself direct Cas9 to the target (compare the upper and lower panels of Figure 2).

[0071] Mali et al. (Science. 2013 Feb 15; 339(6121):823-826.), fully incorporated herein by reference, subsequently adapted the system for use in mammalian cells by providing (i) an ORF encoding codon-optimized Cas9 (e.g., class II, type II Cas enzyme) under a suitable promoter having a C-terminal nuclear localization sequence (e.g., SV40 NLS) and a suitable polyadenylation signal (e.g., TK pA signal), and (ii) an ORF encoding sgRNA under a suitable polymerase III promoter (e.g., U6 promoter) (having a 5' sequence beginning with G, followed by a 3' tracr binding sequence, a linker, and a 20nt complementary target nucleic acid sequence linked to the tracrRNA sequence).

[0072] Transposons are mobile elements that can move within the genome. Such transposons have evolved to limit their adverse effects on the host. Various regulatory mechanisms maintain transposons at low frequencies and are sometimes used to coordinate them with various cellular processes. Some prokaryotic transposons can mobilize functions that benefit the host or otherwise help maintain the element. Certain transposons may have further evolved mechanisms that tightly control the selection of target sites, the most notable example of which is the Tn7 family.

[0073] Transposons Tn7 and similar elements not only encode other adaptive functions in the natural environment, but may also be reservoirs of antibiotic resistance and pathogenic function in the clinical environment. For example, while the Tn7 system has evolved mechanisms to almost completely avoid integration into important host genes, it is also possible to maximize the dispersion of elements by recognizing mobile plasmids and bacteriophages that can transfer Tn7 between host bacteria.

[0074] Tn7 and Tn7-like elements can control the location and timing of their insertion, possessing one pathway that induces insertion into a single conserved site within the bacterial genome, and a second pathway that appears to be better suited to maximizing targeting to mobile plasmids capable of transporting the element between bacteria (see Figure 3). The association between Tn7-like transposons and the CRISPR-Cas system suggests that the transposons may hijack CRISPR effectors to generate R-loops at target sites, thereby facilitating transposon dispersal via plasmids and phages.

[0075] MG64 series

[0076] In one embodiment, the present disclosure provides a system for transposing a cargo nucleotide sequence to a target nucleic acid site. The system may include a first double-stranded nucleic acid containing a cargo nucleotide sequence. This cargo nucleotide sequence may be configured to interact with a Tn7-type transposase complex. The system may also include a Cas effector complex. The Cas effector complex may include a class II, type V Cas effector and an engineered guide polynucleotide configured to hybridize to the target nucleotide sequence. The system may also include a Tn7-type transposase complex configured to bind to the Cas effector complex, the Tn7-type transposase complex containing a TnsB subunit.

[0077] In some cases, the cargo nucleotide sequence is adjacent to the left transposase recognition sequence. In some cases, the cargo nucleotide sequence is adjacent to the right transposase recognition sequence. In some cases, the cargo nucleotide sequence is adjacent to both the left and right transposase recognition sequences. In some cases, the system further includes a second double-stranded nucleic acid containing the target nucleic acid site. In some cases, the system further includes a PAM sequence that fits the Cas effector complex adjacent to the target nucleic acid site. In some cases, the PAM sequence is located at 3' of the target nucleic acid site.

[0078] In some cases, the manipulated guide polynucleotide is configured to bind to a class II, type V Cas effector. In some cases, the class II, type V Cas effector is a class II, type VK effector. In some cases, a Class II, Type V Cas effector contains a polypeptide containing a sequence having at least approximately 20%, at least approximately 25%, at least approximately 30%, at least approximately 35%, at least approximately 40%, at least approximately 45%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, or at least approximately 99% identity with SEQ ID NOs. In some cases, a Class II, Type V Cas effector contains a polypeptide with a sequence substantially identical to SEQ ID NOs. 1, 12, 16, 20-30, 64, or 80-85. In some cases, a TnsB subunit contains a polypeptide with a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NOs. 2, 13, 17, or 65, or their variants. In some cases, the TnsB subunit contains a polypeptide having a substantially identical sequence to sequence numbers 2, 13, 17, or 65.

[0079] In some cases, the Tn7 type transposase complex contains at least one polypeptide containing a sequence having at least approximately 20%, at least approximately 25%, at least approximately 30%, at least approximately 35%, at least approximately 40%, at least approximately 45%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, or at least approximately 99% identity with any one of SEQ ID NOs. 3-4, 14-15, 18-19, or 66-67 or their variants. In some cases, the recombinase complex contains at least one polypeptide having a substantially identical sequence to one of SEQ ID NOs: 3-4, 14-15, 18-19, or 66-67. In some cases, the Tn7 type transposase complex contains at least two polypeptides containing sequences that have at least approximately 20%, at least approximately 25%, at least approximately 30%, at least approximately 35%, at least approximately 40%, at least approximately 45%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, or at least approximately 99% identity with any one of SEQ ID NOs. 3-4, 14-15, 18-19, or 66-67 or their variants. In some cases, the Tn7 type transposase complex contains at least two polypeptides that have substantially identical sequences to one of sequence numbers 3-4, 14-15, 18-19, or 66-67.

[0080] In some cases, the manipulated guide polynucleotide contains a sequence of at least 46 to 80 consecutive nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs. 5-6, 32-33, 94-95, or 104-105 or their variants. In some cases, the manipulated guide polynucleotide contains a sequence of at least approximately 46–80 consecutive nucleotides substantially identical to one of sequence numbers 5–6, 32–33, 94–95, or 104–105.

[0081] In some cases, the left-hand recombinase sequence contains a sequence that is at least approximately 20%, at least approximately 25%, at least approximately 30%, at least approximately 35%, at least approximately 40%, at least approximately 45%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, or at least approximately 99% identical to sequence numbers 9, 11, 36–38, 76, or 78.

[0082] In some cases, the recombinase sequence on the right contains a sequence that is at least approximately 20%, at least approximately 25%, at least approximately 30%, at least approximately 35%, at least approximately 40%, at least approximately 45%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, or at least approximately 99% identical to sequence number 8, 10, 39–44, 77, 79, or 93.

[0083] In some cases, class II, type V Cas effectors and Tn7 type transposase complexes are encoded by polynucleotide sequences containing less than approximately 20 kilobases, less than approximately 15 kilobases, less than approximately 10 kilobases, or less than approximately 5 kilobases.

[0084] In one embodiment, the present disclosure provides a method for transposing a cargo nucleotide sequence to a target nucleic acid site, the method comprising the steps of expressing the system described herein in a cell or introducing the system described herein into a cell.

[0085] In one embodiment, the Disclosure provides a method for transposing a cargo nucleotide sequence to a target nucleic acid site containing a target nucleotide sequence, the method comprising contacting a first double-stranded nucleic acid containing the cargo nucleotide sequence with a Cas effector complex comprising a class II, type V Cas effector and at least one manipulated guide polynucleotide configured to hybridize to the target nucleotide sequence. The method may also comprise contacting a first double-stranded nucleic acid containing the cargo nucleotide sequence with a Tn7 type transposase complex configured to bind to the Cas effector complex, the Tn7 type transposase complex comprising a TnsB subunit. The method may also comprise contacting a first double-stranded nucleic acid containing the cargo nucleotide sequence with a second double-stranded nucleic acid containing a target nucleic acid site.

[0086] In some cases, the cargo nucleotide sequence is adjacent to the left transposase recognition sequence. In some cases, the cargo nucleotide sequence is adjacent to the right transposase recognition sequence. In some cases, the cargo nucleotide sequence is adjacent to both the left and right transposase recognition sequences. In some cases, the method further includes a PAM sequence that fits the Cas effector complex adjacent to the target nucleic acid site. In some cases, the PAM sequence is located at 3' of the target nucleic acid site.

[0087] In some cases, the manipulated guide polynucleotide is configured to bind to a class II, type V Cas effector. In some cases, the class II, type V Cas effector contains a polypeptide containing a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NOs. In some cases, Class II, Type V Cas effectors contain polypeptides with substantially identical sequences to sequence numbers 1, 12, 16, 20-30, 64, or 80-85.

[0088] In some cases, the TnsB subunit contains a polypeptide having a sequence that is at least approximately 20%, at least approximately 25%, at least approximately 30%, at least approximately 35%, at least approximately 40%, at least approximately 45%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, or at least approximately 99% identical to sequence number 2, 13, 17, or 65. In some cases, the TnsA subunit contains a polypeptide having a sequence that is substantially identical to sequence number 2, 13, 17, or 65.

[0089] In some cases, the Tn7 type transposase complex contains at least one polypeptide containing a sequence having at least approximately 20%, at least approximately 25%, at least approximately 30%, at least approximately 35%, at least approximately 40%, at least approximately 45%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, or at least approximately 99% identity with any one of SEQ ID NOs. 3-4, 14-15, 18-19, or 66-67 or their variants. In some cases, the recombinase complex contains at least one polypeptide having a substantially identical sequence to one of SEQ ID NOs: 3-4, 14-15, 18-19, or 66-67. In some cases, the Tn7 type transposase complex contains at least two polypeptides containing sequences that have at least approximately 20%, at least approximately 25%, at least approximately 30%, at least approximately 35%, at least approximately 40%, at least approximately 45%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, or at least approximately 99% identity with any one of SEQ ID NOs. 3-4, 14-15, 18-19, or 66-67 or their variants. In some cases, the Tn7 type transposase complex contains at least two polypeptides that have substantially identical sequences to one of sequence numbers 3-4, 14-15, 18-19, or 66-67.

[0090] In some cases, the manipulated guide polynucleotide contains a sequence of at least 46 to 80 consecutive nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs. 5-6, 32-33, 94-95, or 104-105 or their variants. In some cases, the manipulated guide polynucleotide contains a sequence of at least approximately 46–80 consecutive nucleotides substantially identical to one of sequence numbers 5–6, 32–33, 94–95, or 104–105.

[0091] In some cases, the left-hand recombinase sequence contains a sequence that is at least approximately 20%, at least approximately 25%, at least approximately 30%, at least approximately 35%, at least approximately 40%, at least approximately 45%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, or at least approximately 99% identical to sequence numbers 9, 11, 36–38, 76, or 78. In some cases, the recombinase sequence on the right contains a sequence that is at least approximately 20%, at least approximately 25%, at least approximately 30%, at least approximately 35%, at least approximately 40%, at least approximately 45%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, or at least approximately 99% identical to sequence number 8, 10, 39–44, 77, 79, or 93.

[0092] In some cases, class II, type V Cas effectors and Tn7 type transposase complexes are encoded by polynucleotide sequences containing less than approximately 20 kilobases, less than approximately 15 kilobases, less than approximately 10 kilobases, or less than approximately 5 kilobases.

[0093] In accordance with the IUPAC Agreement, the following abbreviations will be used throughout the examples. A = Adenine C = Cytosine G = Guanine T = Chimin R = adenine or guanine Y = Cytosine or Thymine S = Guanine or Cytosine W = Adenine or Thymine K = guanine or thymine M = adenine or cytosine B = C, G, or T D = A, G, or T H = A, C, or T V = A, C, or G [Examples]

[0094] Example 1 - (General Protocol) Identification / Confirmation of PAM Sequences for the Systems Described herein The putative endonuclease was expressed in an E. coli lysate-based expression system (myTXTL, Arbor Biosciences). The PAM sequence was determined by sequencing a plasmid containing a randomly generated potential PAM sequence that could be cleaved by the putative nuclease. In this system, the E. coli codon-optimized nucleotide sequence encoding the putative nuclease was transcribed and translated in vitro from a PCR fragment under the control of the T7 promoter. A second PCR fragment with the T7 promoter and a minimal CRISPR array consisting of a subsequent repeat-spacer-repeat sequence was also transcribed in the same reaction. Successful expression of the endonuclease and repeat-spacer-repeat sequence in the TXTL system, followed by CRISPR array processing, yielded an active in vitro CRISPR nuclease complex.

[0095] A library of target plasmids containing spacer sequences matching a minimal array spacer sequence preceded by an 8N mixed group (potential PAM sequence) was incubated at the yield of a TXTL reaction. After 1–3 hours, the reaction was stopped, and DNA was recovered using a DNA cleanup kit, e.g., Zymo DCC, AMPure XP beads, QiaQuick. The adapter sequence was blunt-end ligated to DNA containing the active PAM sequence cleaved by an endonuclease, while uncleaved DNA was inaccessible for ligation. Subsequently, DNA segments containing the active PAM sequence were amplified by PCR using primers specific to the library and the adapter sequence. The PCR amplification products were degraded on a gel, and amplicons corresponding to the cleavage events were identified. The amplified segments from the cleavage reaction were also used as templates for NGS library preparation or as substrates for Sanger sequencing. Sequence analysis of this resulting library, a subset of the starting 8N library, revealed sequences with PAM activity compatible with the CRISPR complex. For PAM testing using processed RNA constructs, the same procedure was repeated, except that in vitro transcribed RNA was added along with the plasmid library, and a minimal CRISPR array template was omitted.

[0096] Analysis of the intergenetic regions surrounding the Cas effector and CRISPR array identified potential anti-repeat sequences corresponding to the double-stranded tracrRNA sequence. The tracrRNA and crRNA repeats were folded and trimmed, and a tetraloop sequence of GAAA was added to maintain the stem-loop region of the crRNA-tracrRNA complex.

[0097] Example 2a - In vitro targeted integrase activity Integrase activity was assayed preferentially using previously identified PAMs, but may instead be performed with lower efficiency using PAM library substrates. One configuration of components for the in vitro test included (1) an expression plasmid with an effector (or multiple effectors) under the T7 promoter, (2) an expression plasmid with a transposase gene under the T7 promoter, sgRNA or crRNA, and tracrRNA, (3) a target plasmid containing a spacer site and appropriate PAMs, and (4) a donor plasmid containing the left-end (LE) and right-end (RE) DNA sequences necessary for transposition around a cargo gene (e.g., a select marker such as the Tet resistance gene), with three plasmids in addition to the one containing the donor sequence. Effector and transposase genes were expressed using an in vitro transcription / translation (TXTL) system (e.g., an E. coli lysate-based or reticulocyte extract-based system). After expression, RNA, target DNA, and donor DNA were added and incubated to induce transposition. Transposition was detected by PCR across the transposase junction with one primer on the target DNA and one on the donor DNA. The resulting PCR products were sequenced by NGS to determine the precise insertion topology to the sgRNA / crRNA target site. The primers were positioned downstream to accommodate and detect various insertion sites. The primers were designed so that integration could be detected in any direction of the cargo and on either side of the spacer, as the direction of integration was initially unknown.

[0098] The integration efficiency was measured by quantitative PCR (qPCR) of the experimental output of target DNA containing the integrated cargo, and normalized to the amount of unmodified target DNA, also measured by qPCR.

[0099] This assay can be performed using purified protein components rather than lysate-based expression. In this case, the protein was expressed in E. coli protease-deficient strain B under a T7-inducible promoter, the cells were lysed using sonication, and the His-tagged protein of interest was purified using HisTrap FF (GE Lifescience) Ni-NTA affinity chromatography on AKTA Avant FPLC (GE Lifescience). Purity was determined by SDS-PAGE and densitometry using ImageLab software (Bio-Rad) for protein bands separated on InstantBlue Ultrafast (Sigma-Aldrich) Coomassie-stained acrylamide gel (Bio-Rad). The protein was desalted in a storage buffer consisting of 50 mM Tris-HCl, 300 mM NaCl, 1 mM TCEP, and 5% glycerol pH 7.5 (or other buffer determined to obtain maximum stability) and stored at -80°C. After purification, the effector and transposase were added to the sgRNA, target DNA, and donor DNA as described above in a reaction buffer consisting of 26 mM HEPES pH 7.5, 4.2 mM TRIS pH 8, 50 μg / mL BSA, 2 mM ATP, 2.1 mM DTT, 0.05 mM EDTA, 0.2 mM MgCl2, 28 mM NaCl, 21 mM KCl, and 1.35% glycerol (final pH 7.5) supplemented with 15 mM Mg(OAc)2.

[0100] Example 2b-in vitro activity Targeted nucleases

[0101] In situ expression and protein sequence analysis suggested that several RNA guide effectors were active nucleases. They contained predicted endonuclease-associated domains (corresponding to RuvC and HNH_endonuclease domains) and / or predicted HNH and RuvC catalytic residues.

[0102] The activity of candidate proteins was tested with engineered single guide RNA sequences using the myTXTL system and in vitro transcription RNA. Active proteins that successfully cleaved the library spawned a band around 170 bp in the gel.

[0103] DNA integration and transposition

[0104] A transposon is predicted to be active when the encoding genomic sequence contains one or more protein sequences with transposase and / or integrase function within the left and right ends of the transposon. The Tn7 transposon consists of the catalytic transposase TnsB as defined herein, but may also contain TnsA, TnsC, TnsD, TnsE, TniQ, and / or other transposases or integrases. The transposon terminus consists of a predicted transposase binding site, which contains 15 bp–150 bp long direct and / or reverse repeats adjacent to the transposase protein and other “cargo” genes. Protein sequence analysis shows that the transposase contains an integrase domain, a transposase domain, and / or transposase catalytic residues, suggesting that they are active (e.g., Figure 4A).

[0105] Integration of targeted DNA

[0106] A putative CRISPR-associated transposon (CAST) contains a CRISPR nuclease or effector targeting DNA and / or RNA, along with a protein possessing predicted transposase function, located near a CRISPR array. In some systems, the nuclease is predicted to be active based on the presence of an endonuclease-associated catalytic domain and / or catalytic residues.

[0107] In some systems, the effector is predicted to be homologous to known CRISPR effector proteins but inactive based on the absence of the endonuclease domain and / or catalytic residues. The transposase is predicted to be associated with the effector when the CRISPR locus (inactive CRISPR nuclease and array) and the transposase protein are located within the left and right ends of the predicted transposon (Figure 4A). In this case, the effector is predicted to direct DNA integration to a specific genomic location based on guide RNA.

[0108] CAST activity was tested using five types of components: (1) Cas effector protein expressed by myTXTL or PURExpress, (2) target DNA fragment or plasmid containing the target sequence and PAM corresponding to the Cas enzyme, (3) donor DNA fragment containing a DNA marker or fragment with LE and RE on both sides of the transposase system in the DNA fragment or plasmid, (4) any combination of transposase proteins expressed using myTXTL or PURExpress, and (5) engineered in vitro transcribed single guide RNA sequences. Active systems with successfully transposed donor fragments were assayed by PCR amplification of the donor-target junction.

[0109] After the rearrangement reaction, PCR amplification of the junction site revealed appropriate donor-target formation, demonstrating that the rearrangement reaction is sg-dependent (Figure 6). PCR amplification of reactions #3 and #4 showed that both orientations of the donor to the target occurred: one where LE is close to PAM, and another where RE is close to PAM. While both rearrangement orientations occurred, the orientation where LE is close to PAM was preferred for donor integration into the target, as indicated by strong bands present in reactions #4 and #5.

[0110] Sanger sequencing was performed on the products of the preferred orientations described above. Among the incorporations that occurred with LE close to PAM, there was a clear degradation of the sequencing chromatogram signal from either the forward or reverse direction across the target / donor junction. This indicated that among the products oriented with LE close to PAM, the incorporation occurred within the nucleotide range, and the primary product of the product with LE close to PAM was a 61 bp incorporation from PAM (Figure 7A). Donor-derived sequencing across the donor-target junction defined the essential outer boundary configuration of the LE and RE sequences (Figures 7A and B). Further investigation of the LE and RE domains determined the internal limits of the LE and RE sequences essential for the transposition. Sequencing of RE in products with LE close to PAM revealed a 3 bp duplication downstream of the donor RE (Figure 7B). This is partly due to a Tn7 transposase incorporation event that cleaved and ligated the donor fragment at a staggered cut site. The 3bp overlap is smaller than the 5bp overlap expected from other Tn7 transposases.

[0111] Sanger sequencing of PCR amplification products in an 8N library of the target plasmid revealed the PAM preference for the MG64-1 effector as nGTn / nGTt at the 5' end of the spacer (Figure 7C). NGS analysis of the PAM library target confirmed the preference for the nGTn motif at the 5' end.

[0112] Example 3 - Predicted RNA folding The predicted RNA folding of a single active RNA sequence was calculated at 37° using the Andronescu 2007 method. All hairpin loop secondary structures were individually deleted from the structure and repeatedly compiled into smaller single guides. In a second approach, the tracrRNA of MG64-1 was aligned with the tracrRNA of a known type Vk, and the unique insertion region was mutated from the single guide and minimized to 57 bases. Figure 12A shows the predicted structure of MG64-1 sgRNA. Figure 12B shows the predicted structure of MG64-3 sgRNA. Figure 12C shows the predicted structure of MG64-5 sgRNA. The color of the bases corresponds to the probability of base pairing, with red representing a high probability and blue representing a low probability.

[0113] Example 4 - Verification of transposon ends by gel shift Transposon ends were tested for TnsB binding via electrophoretic mobility shift assay (EMSA). In this case, potential LE or RE were synthesized as DNA fragments (100–500 bp) and end-labeled with FAM via PCR using FAM-labeled primers. TnsB proteins were synthesized in an in vitro transcription / translation system (e.g., PURExpress). After synthesis, 1 μL of TnsB protein was added to 50 nM labeled RE or LE in a 10 μL reaction in binding buffer (20 mM HEPES pH 7.5, 2.5 mM Tris pH 7.5, 10 mM NaCl, 0.0625 mM EDTA, 5 mM TCEP, 0.005% BSA, 1 ug / mL poly(dI-dC), and 5% glycerol). After incubation at 30°C for 40 minutes, 2 μL of 6X loading buffer (60 mM KCl, 10 mM Tris, pH 7.6, 50% glycerol) was added. The binding reaction was isolated and visualized on a 5% TBE gel. A shift in LE or RE in the presence of TnsB was attributed to successful binding and indicated transposase activity (Figure 24).

[0114] Example 5 - Integrase activity in Escherichia coli Because E. coli lacks the ability to efficiently repair double-strand DNA breaks in its genome, transformation of E. coli with drugs that can induce double-strand breaks in the E. coli genome leads to cell death. This phenomenon was used to test endonuclease or effector-assisted integrase activity in E. coli by recombinantly expressing either the endonuclease or effector-assisted integrase with guide RNA (determined, e.g., as in Example 3) in target strains having spacer / target and PAM sequences integrated into their genomic DNA.

[0115] Subsequently, the manipulated strains were transformed with plasmids containing a nuclease or effector with a single guide RNA, plasmids expressing integrase and accessory genes, and plasmids containing a temperature-sensitive origin of replication with selectable markers flanked by left-end (LE) and right-end (RE) transposon motifs for integration. Transformants induced for the expression of these genes were screened for introduction of markers into genomic targets by selection at limiting temperatures for plasmid replication, and integration of the markers into the genome was confirmed by PCR.

[0116] Off-target insertions were screened using an unbiased approach. Briefly, purified gDNA was fragmented by Tn5 transposase or shearing, and then the target DNA was PCR-amplified using primers specific to the ligated adapter and selectable markers. The amplicons were then prepared for NGS sequencing. The resulting sequences were analyzed by trimming the transposon sequences, mapping the flanking sequences to the genome to determine insertion sites, and determining the off-target insertion rate.

[0117] Example 6 - Colony PCR screening of transposase activity To test nuclease or effector-assisted integrase activity in bacterial cells, strain MGB0032 was constructed from BL21(DE3) E. coli cells engineered to contain target and corresponding MG64_1-specific PAM sequences. Subsequently, MGB0032 E. coli cells were transformed with pJL56 (a plasmid expressing a combination of MG64_1 effector and helper, ampicillin-resistant) and pTCM64_1sg, a chloramphenicol-resistant plasmid expressing an engineered single guide RNA sequence of the target in question, driven by the T7 promoter.

[0118] Next, MGB0032 cultures containing both plasmids were grown to saturation, diluted at least 1:10 in growth cultures containing appropriate antibiotics, and incubated at 37°C until the OD was approximately 1. Cells from this growth stage were electrocompetent and transformed and incorporated with streamlined 64_1 pDonor, a plasmid containing tetracycline resistance markers flanked by the left-end (LE) and right-end (RE) transposon motifs. Electroporated cells were harvested in LB medium for 2 hours in or without a final concentration of 100 μM IPTG, then plated in LB-agar-ampicillin-chloramphenicol-tetracycline and incubated at 37°C for 4 days. Each resulting CFU was sampled using a sterile toothpick and mixed with water. To this solution, we added the Q5 high-fidelity PCR master mix (New England Biolabs) and primers LA155 (5'-GCTCTTCCGATCTNNNNNGATGAGCGCATTGTTAGATTTCAT-3') and oJL50 (5'-AAACCGACATCGCAGGCTTC-3'). These primers were positioned adjacent to the predicted insertion junction. The predicted product size was 609 bp. The DNA-amplified PCR product was visualized on a 2% agarose gel. Sanger sequencing of the PCR product confirmed the transposition event.

[0119] Example 7 - Intracellular Expression / In vitro Assay To test the functionality of NLS constructs in physiologically relevant environments, constructs cloned with active NLS-tagged CAST components were incorporated into K562 cells using lentiviral transfection. Briefly, constructs cloned into lentiviral transfection plasmids were transfected into 293T cells containing envelope plasmids and packaging plasmids. After 72 hours of incubation, the virus-containing supernatant was collected from the culture medium. Next, the virus-containing medium was incubated with 8 μg / mL polyblen in the K562 cell line for 72 hours. The transfected cells were then selected for large-scale incorporation using 1 μg / mL puromycin for 4 days. The selected cell lines were harvested at the end of the 4 days and lysed separately to obtain nuclear and cytoplasmic fractions. The fractions were then tested for translocation ability using complementary sets of components expressed in vitro.

[0120] Ten million cells were centrifuged and washed once with 1×PBS pH 7.4. The supernatant was completely aspirated into the cell pellet and flash-frozen at -80°C for 16 hours. After thawing on ice, the cell pellet size was measured by mass, and proteins from the cell fraction were spontaneously extracted using appropriate extraction volumes of cell fraction and nuclear extraction reagent (NE-PER). Briefly, the cytoplasmic extraction reagent was used in a ratio of 1:10 cell mass to extraction reagent. The cell suspension was mixed by vortexing and dissolved with a nonionic washing agent. The cells were then centrifuged at 16,000×g at 4°C for 5 minutes. Next, the cytoplasmic extract supernatant was decanted and stored for in vitro testing. Then, the nuclear extraction reagent was added in a ratio of 1:2 to the original cell mass and incubated on ice for 1 hour with intermittent vortexing. Subsequently, the nuclear suspension was centrifuged at 16,000 × g for 10 minutes at 4°C, and the supernatant nuclear extract was decanted and tested for in vitro translocation activity. Using 4 μL of each cell and nuclear extract from each condition, in vitro translocation reactions were performed with complementary sets of in vitro expressed proteins, donor DNA, pTarget, and buffer. Evidence of translocation activity was assayed by PCR amplification of the donor-target junction.

[0121] Example 8 - Activity in mammalian cells (predictive) To demonstrate targeting and cleavage activity in mammalian cells, nuclear localization sequences are fused to the C-terminus of nuclease or effector proteins, and the fusion proteins are purified with the integrase proteins. A single guide RNA targeting the genomic locus of interest is synthesized and incubated with the nuclease / effector protein to form a ribonucleoprotein complex. Cells are transfected with plasmids containing selectable neomycin resistance markers (NeoR) or fluorescent markers adjacent to the left-end (LE) and right-end (RE) motifs, harvested for 4–6 hours, and then electroporated with nuclease RNP and integrase proteins. Plasmid integration into the genome is quantified by counting G418-resistant colonies or by cytometry of fluorescence-activated cells. Genomic DNA is extracted 72 hours after electroporation and used for NGS library preparation. Off-target frequencies are assayed by fragmenting the genome and preparing amplicons of transposon markers and adjacent DNA for NGS library preparation. To test the activity of each targeting system, at least 40 different target sites are selected.

[0122] Example 9 - Activity of the targeted nuclease In situ expression and protein sequence analysis suggested that several RNA guide effectors are active nucleases. These contain predicted endonuclease-associated domains (corresponding to RuvC and HNH_endonuclease domains) and predicted HNH and RuvC catalytic residues (Figure 4A).

[0123] The activity of candidate proteins was tested using the myTXTL system and in vitro transcription RNA with manipulated single guide RNA sequences. Successfully cleaved active proteins in the library yielded a band of approximately 170 bp in the gel.

[0124] Example 10 - Identification of transposons A transposon is predicted to be active when it contains one or more protein sequences having transposase and / or integrase function between its left and right ends. The Tn7 transposon consists of the catalytic transposase TnsB as defined herein, but may also contain TnsA, TnsC, TnsD, TnsE, TniQ, and / or other transposases or integrases. The transposon terminus consists of a predicted transposase binding site, which contains 15 bp–150 bp long direct and / or reverse repeats adjacent to the transposase protein and other “cargo” genes. Protein sequence analysis shows that the transposase contains an integrase domain, a transposase domain, and / or transposase catalytic residues, suggesting that they are active (e.g., Figure 4A and Figure 5A).

[0125] Example 11 - Identification of CRISPR-related transposons A putative CRISPR-associated transposon (CAST) contains a CRISPR effector targeting DNA and / or RNA, and a protein with predicted transposase function, located near the CRISPR array. In some systems, the effector is predicted to have nuclease activity based on the presence of an endonuclease-related catalytic domain and / or catalytic residues (e.g., Figure 4A). The transposase was predicted to be associated with an active nuclease when the CRISPR locus (CRISPR nuclease and array) and the transposase protein are located between the left and right ends of the predicted transposon (e.g., Figure 4B and C). In this case, the effector was predicted to direct DNA integration to a specific genomic location based on guide RNA.

[0126] In some systems, the effector was predicted to be homologous to known CRISPR effector proteins but inactive based on the absence of the endonuclease domain and / or catalytic residues (Figure 5A). The transposase was predicted to be associated with the effector when the CRISPR locus (inactive CRISPR nuclease and array) and the transposase protein were located within the left and right ends of the predicted transposon (Figures 5A and 5B).

[0127] Example 12 - Discovery of CAST CRISPR-related transposons (CASTs) are a system of transposons that have evolved to interact with the CRISPR system to facilitate the targeted integration of DNA cargo.

[0128] CAST is a genomic sequence encoding one or more protein sequences involved in DNA transposition within the left and right ends of a transposon signature. A Tn7 transposon, as defined here, consists of the catalytic transposase TnsB, but may also include the catalytic transposase TnsA, the loader protein TnsC or TniB, and the target recognition proteins TnsD, TnsE, TniQ, and / or other transposon-related components. The transposon terminus consists of a predicted transposase binding site, which includes direct and / or reverse repeats of 15 bp–150 bp in length, adjacent to the transposon mechanism and other “cargo” genes.

[0129] In addition, CAST further encodes DNA and / or RNA that target CRISPR nucleases or effectors in the vicinity of the CRISPR array. In some systems, the effectors were predicted to be active nucleases based on the presence of an endonuclease-associated catalytic domain and / or catalytic residues. In some systems, the effectors were predicted to be inactive based on the absence of an endonuclease domain and / or catalytic residues, although they had sequence similarity to known CRISPR effector proteins. Transposons were predicted to be associated with effectors if the CRISPR locus and transposon-associated proteins were located within the left and right ends of the predicted transposon. In this case, the effectors were predicted to direct DNA integration to a specific genomic location based on guide RNA.

[0130] Example 13 - Class II Cas12K CAST The Cas12k CAST system encodes a nuclease-deficient CRISPR Cas12k effector, a CRISPR array, tracrRNA, and a Tn7-like transposition protein. The Cas12k effector is phylogenetically diverse, and several features confirming its association with CAST have been identified (Figure 8). For example, the left end of the transposon was identified downstream of the MG64-3 CRISPR locus, as indicated by the terminal reverse repeat sequence and self-congruent spacer sequence (Figure 11A). The Cas12k CAST CRISPR repeat (crRNA) contains the conserved motif 5'-GNNGGNNTGAAAG-3' (Figure 9). Short repeat-antirepeats (RARs) within crRNA motifs aligned with different regions of tracrRNA (Figures 9 and 10), and the RAR motifs appeared to define the start and end of tracrRNA (for example, for MG64-1, the 5' end of tracrRNA contained RAR1 (TTTC) and the 3' end contained RAR2 (CCNNC) (Figure 10A)).

[0131] Example 14 - Transposon End Prediction The transposon terminus was predicted from the intergeneric region adjacent to the effector and the transposon mechanism. For example, for Cas12k CAST, the intergeneric region directly located upstream of TnsB and directly downstream of the CRISPR locus was predicted to contain the left and right ends (LE and RE) of the Tn7 transposon.

[0132] Direct and reverse repeats (DR / IR) of approximately 12 bp, with up to two mismatches, were predicted on the contig. In addition, short (approximately 10-20 bp) DR / IRs adjacent to the CAST transposon were discovered using the Dotplot algorithm. Matching DR / IRs located in intergeneric regions adjacent to the CAST effector and the transposon gene are predicted to encode the transposon binding site. The transposon terminal boundary was defined by aligning the LE and RE extracted from the intergeneric region encoding the putative transposon binding site. The LE and RE terminals of the putative transposon are: a) regions located within 400 bp upstream and downstream of the first and last predicted transposon-coding genes, b) regions sharing multiple short reverse repeats, and c) regions sharing more than 65% of nucleotide IDs.

[0133] Example 15 - Single Guide Design Analysis of the intergenetic regions surrounding the Cas effector and CRISPR array identified potential antirepeat sequences and a conserved "CYCC(n6)GGRG" stem-loop structure adjacent to the antirepeats corresponding to the double-stranded tracrRNA sequence (Figure 11B). The tracrRNA and crRNA repeats were folded and trimmed, and a tetraloop sequence of GAAA was added to maintain the stem-loop region of the crRNA-tracrRNA complementary sequence.

[0134] Example 16 - In vitro integration activity using the targeted nuclease In situ expression and protein sequence analysis suggested that several RNA guide effectors are active nucleases. They contain predicted endonuclease-associated domains (matching RuvC and HNH_endonuclease domains) and / or predicted HNH and RuvC catalytic residues. Candidate activity was tested with engineered single guide RNA sequences using the myTXTL system and in vitro transcription RNA. Active proteins successfully cleaved from the library yielded a band per 170 bp in the gel.

[0135] Example 17 - Programmable DNA Integration CAST activity was tested using five components: (1) Cas effector protein expressed by myTXTL or PURExpress (SEQ ID NO: 1), (2) target DNA fragment or plasmid containing the target sequence and PAM corresponding to the Cas enzyme (SEQ ID NO: 31), (3) donor DNA fragment containing a marker or DNA fragment with LE and RE on both sides of the transposase system in the DNA fragment or plasmid (SEQ ID NOs: 8-11), (4) any combination of transposase proteins expressed using myTXTL or PURExpress (SEQ ID NOs: 2-4), and (5) a manipulated in vitro transcribed single guide RNA sequence (SEQ ID NO: 5). The active system with successfully transposed donor fragments was assayed by PCR amplification of the donor-target junction.

[0136] After the rearrangement reaction, PCR amplification of the junction site revealed appropriate donor-target formation, demonstrating that the rearrangement reaction is sg-dependent (Figure 9). PCR amplification of reactions #3 and #4 showed that both orientations of the donor to the target occurred: one where LE is close to PAM, and another where RE is close to PAM. While both rearrangement orientations occurred, a preference was observed for donor integration into the target when LE is close to PAM, as represented by strong bands present in reactions #4 and #5.

[0137] Sanger sequencing was performed on the products of the preferred orientations described above. Among the integrations that occurred with LEs close to PAM, there was a clear degradation of the sequencing chromatogram signal from either the forward or reverse direction across the target / donor junction. This indicated that among the products oriented with LEs close to PAM, the integration occurred within the nucleotide range, and the primary product of the product with LEs close to PAM was a 61 bp integration from PAM (Figure 10A). Donor-derived sequencing across the donor-target junction defined the construction of the essential outer boundary of the LE and RE sequences (Figures 10A, 10B). Sequencing of the RE in products with LEs close to PAM revealed a 3 bp overlap downstream of the donor RE (Figure 10B). This is partly due to a Tn7 transposase integration event that cleaved and ligated the donor fragment at a staggered cut site. The 3 bp overlap is smaller than the 5 bp overlap expected from other Tn7 transposases.

[0138] Furthermore, Sanger sequencing of PCR amplification products in an 8N library of the target plasmid demonstrated PAM preference for the MG64-1 effector as nGTn / nGTt at the 5' end of the spacer (Figure 10C). NGS analysis of the PAM library target confirmed the preference for the nGTn motif at the 5' end.

[0139] Further development of single-guide assays using a new sgRNA scaffold confirmed the activity of MG64-1 (Figure 13).

[0140] Example 18 - Determination of the Integration Window The PCR junctions of the amplified PAM were indexed against an NGS library and sequenced using MiSeq with a V2 300 read kit. Reads were mapped and quantified using CRISPResso, which uses the amplicon sequence of the estimated transposition sequence (guideseq=LE or RE 3' end 20bp, window center=0, window size=20) with an integration distance of 60bp from the PAM. The indel histogram was normalized for all detected indel reads, and the frequencies were plotted against the 60bp reference sequence (Figure 14).

[0141] Both PCR reaction 5 (LE proximal to PAM, upper panel of Figure 14) and PCR 4 (RE distal to PAM, lower panel of Figure 14) were plotted for MG64-1 on the sequence and distance from PAM. Integration window analysis showed that 95% of integrations occurring at the spacer PAM site were within a 10 bp window, 58–68 nucleotides away from PAM. The difference in integration distance between distal and proximal frequencies reflected overlapping integration sites, i.e., 3–5 base pair overlaps, resulting from shifts in transposase nuclease activity during integration.

[0142] Example 19 - Colony PCR screening of transposase activity Transcriptional activity was assayed via colony PCR screening. After transformation with the pDonor plasmid, E. coli were seeded on LB-agar containing ampicillin, chloramphenicol, and tetracycline. Selected CFUs were added to a solution containing PCR reagents and primers adjacent to the selected insertion junction. The PCR reaction of the incorporated products was visible on the gel (Figure 15). Sequencing of the selected colony PCR products confirmed that these products exhibited a transposition event when crossing the junction between LE and PAM at the manipulated target site in the lacZ gene.

[0143] Example 20 - Single-Guide Operation The predicted RNA folding of active single RNA sequences was calculated at 37° using the Andronescu 2007 method. All hairpin loop secondary structures were deleted individually from the construct and repeatedly compiled into smaller single guides. Manipulated single guides (esg) 4, 6, 7, 8, and 9 were active for donor transposition (Figure 17 C and D), while manipulated sgRNAs 8 and 9 were weaker single guides and transposed by PCR 5 (Figure 17 D). Manipulated guide 5 was capable of transposition, while manipulated sgRNA 10 transposed weakly by PCR 5 (Figure 17 E and F). Esg17 was a combination of deletions in esg6 and esg7, and esg18 was a combination of esg4 and esg5. Both were able to strongly transpose across both PCR4 and PCR5 (Figure 17G and H), but adding a combination of esg6 and esg18 to create esg19 resulted in a weaker transposition in PCR5, and adding esg7 to esg19 to create esg20 resulted in a junction with very weak transposition relative to PCR5 (Figure 8G and H). In the second approach, the tracrRNA of MG64-1 was aligned with the tracrRNA of a known type Vk, and the unique insertion region was mutated from a single guide. The sgRNA was minimized by truncation of the insertion sequence of the MG64-1 sgRNA (Figure 14). Two subsequent deletions, esg2 and esg3, were also tested (Figure 17A and B), but neither esg2 nor esg3 resulted in a large transposition, and therefore the single guide was minimized by 57 bases.

[0144] Example 21 - LE-RE minimization Sequencing of the target-transition junction helped identify terminal reverse repeats by identifying the outermost sequence from the donor plasmid incorporated into the target reaction. Short repeats contained within the terminals were identified by performing 14 bp repeat analysis with a 10% variability rate, and these minimal terminal truncations were designed to preserve the repeats while deleting the extra sequences. Prediction and cloning were repeated multiple times, and each interaction was tested in vitro with transposition. The initial LE and RE deletions were designed and cloned individually up to 68 bp, 86 bp, and 105 bp for LE, and up to 178 bp, 196 bp, and 242 bp for RE. Since 64-1 RE had further sequences of significant length without repeats, internal deletions of both 50 bp and 81 bp were designed and cloned. The rearrangement in all single deletions was robust to both PCR4 and PCR5 (Figure 18A and B), and the 81bp internal deletion was subsequently continued with a combined deletion in the RE. Trimmed ends of the former 178, 196, and 212bp were cloned over the 81bp internal deletion and tested for rearrangement. The rearrangement was active against all designed constructs. Combined with a 68bp LE, the rearrangement was found to be active up to the 68bp LE region combined with a 96bp RE region.

[0145] Example 22 - Effect of dislocation overhang To test whether extra sequences outside the TnsB binding motif were necessary for transposition, oligos designed for both LE and RE TGTACA motifs were designed and synthesized with extra base pairs of 0, 1, 2, 3, 5, and 10 bp. Using these synthesized oligos, donor PCR fragments with overhangs were generated and their ability to transpose to the target site was tested. Most notably, PCR6, although rarely detected in in vitro reactions (G, lanes 1 and 2 in Figure 18), allowed for efficient incorporation of PCR6 with small 0–3 bp overhangs, reflecting RE proximal to PAM orientation that was not detected with larger adjacent sequences.

[0146] Example 23 - CAST NLS Design Genome editing in eukaryotes for therapeutic purposes heavily relies on the translocation of editing enzymes into the nucleus. A small polypeptide stretch of a large protein signals cellular components to translocate the protein across the nuclear membrane. The placement of NLS tags is crucial, as they must provide translocation functionality while maintaining the function of the protein they are fused to. To test the functional orientation of NLS to each component of the CAST complex, constructs were designed and synthesized in which nucleoplasmin NLS was fused to the N-terminus and SV40 NLS to the C-terminus of each component of MG CAST. Proteins from these constructs were expressed in cell-free in vitro transcription / translation reactions, and in vitro translocation activity was tested using complementary sets of untagged components. The NLS-tagged constructs were evaluated for activity retention by donor-target junction PCR using PCR4 (evaluating distal RE translocation) and the congeneral translocation event, PCR5 (LE to proximal translocation).

[0147] Most components yielded a single NLS orientation that maintained activity. TnsB was the CAST component that was active in both N-terminal and C-terminal NLS by both PCR4 and PCR5 (Figure 19 A, B). TniQ was active with an N-terminal NLS tag (Figure 19 C, D). The Cas12k component was active with a C-terminal tagged NLS (Figure 19 E, F, lanes 5, 6). Further testing of Cas12k with both nucleoplasmin and SV40 NLS tags revealed activity (Figure 19 I, J, lane 4). TnsC was weakly active with an N-terminal NLS (Figure 19 E, F, lane 7), but further investigation of TnsC labeling methods identified novel acting NLS-HA-TnsC and NLS-FLAG-TnsC constructs (Figure 19 G, H, lanes 3 and 7, respectively). The final result was a complete set of NLS-tagged components that were active in vitro in both NLS-TnsB and TnsB-NLS orientations (Lanes A and B, 5 and 6 in Figure 20).

[0148] Example 24 - Design and testing of Cas12k and TniQ protein fusion constructs To simplify the expression of protein components and minimize their delivery into cells, fusion constructs were designed, synthesized, and tested between the Cas12k effector and the TniQ protein. Both orientations of TniQ fused to Cas12k were designed and synthesized, with the C-terminal fusion being Cas-TniQ and the N-terminal fusion being TniQ-Cas. When expressed in vitro and assayed for transposition ability, both constructs showed weak activity to PCR4 (Figure 21A), but the PCR5 junction was firmly formed by the TniQ-Cas fusion protein (Figure 21B). Transposition length was assayed using variable linker domains including the original (20 amino acid linker), 48, 68, 72, and 77 (Figure 21C, D, E, F). Subsequently, NLS tags were ligated to the N-terminus of TniQ and the C-terminus of Cas12k, and the constructs remained active by PCR5 (Figure 20E, F).

[0149] Two other linkers were also employed to fuse the effector gene with the TniQ gene. The self-stopping translation sequence P2A was active in the Cas-NLS-P2A-NLS-TniQ construct (Figure 21, G, H, lanes 6), and the mRNA-based linker of the MCV internal ribosome entry sequence (IRES) enabled independent translation of the two components in the cell (Figure 23, F, G).

[0150] Example 25 - Intracellular expression linked in in vitro translocation test To test the functionality of NLS constructs in physiologically relevant environments, constructs cloned with active NLS-tagged CAST components were incorporated into K562 cells using lentiviral transfection. Briefly, constructs cloned into lentiviral transfection plasmids were transfected into 293T cells containing envelope plasmids and packaging plasmids. After 72 hours of incubation, the virus-containing supernatant was collected from the culture medium. Next, the virus-containing medium was incubated with 8 μg / mL polyblen in the K562 cell line for 72 hours, and the transfected cells were selected for large-scale incorporation using 1 μg / mL puromycin for 4 days. The cell lines to be selected were harvested at the end of the 4 days and lysed separately to obtain nuclear and cytoplasmic fractions. The fractions were then tested for translocation ability using complementary sets of components expressed in vitro.

[0151] Both NLS-TnsB and TnsB-NLS were tested by cell fractionation and in vitro transposition, and when transposition was detected in both cytoplasmic and nuclear fractions, NLS-TniQ had detectable activity in the cytoplasm (Figure 22A, B). Upon expression, both NLS-HA-TnsC and NLS-FLAG-TnsC were active in both cytoplasmic and nuclear fractions (Figure 22D), but PCR4 was formed in the nuclear fraction of both TnsC constructs (Figure 22C).

[0152] When both NLS-TnsB and TnsB-NLS were linked to NLS-FLAG-TnsC using IRES, NLS-TnsB-IRES-NLS-FLAG-TnsC was mostly active in the nuclear fraction, while TnsB-NLS-IRES-NLS-FLAG-TnsC was active in both the cytoplasmic and nuclear fractions. This indicates that NLS-TnsB has an increased ability to traffic to the nucleus (Figure 21E, F).

[0153] Intracellular Cas12k fusions were similarly fractionated and tested for translocation. Cas-NLS-P2A-NLS-TniQ were introduced into cells, fractionated, and tested for intracellular activity in vitro. Cas-NLS-P2A-NLS-TniQ was translocated in the cytoplasm by adding a single guide to the reaction (Figure 23A). The Cas-NLS-P2A-NLS-TniQ construct in the nuclear fraction could be complemented by supplementing with holoCas protein (+sgRNA) or additional TniQ with sgRNA. This indicates that both Cas-NLS and NLS-TniQ can reach the nucleus (Figure 23B, C). The results were similar with the NLS-TniQ-Cas-NLS fusion protein, but required more TniQ supplementation (Figure 23 D, E), while Cas-NLS-IRES-NLS-TniQ required only holo-Cas-NLS supplementation (Figure 23 F, G). Overall, this indicates that all components of CAST were able to be delivered to the nuclear fraction of the cell.

[0154] Example 26 - Verification of transposon ends by gel shift To validate the activity of TnsB against predicted transposon terminal sequences, the LE of MG64-1 was amplified using a FAM-labeled oligonucleotide. The MG64-1 TnsB protein was expressed using a cell-free transcription / translation system and incubated with the LE FAM-labeled product. After incubation for 30 minutes, binding was observed on a native 5% TBE gel (Figure 24). Multiple bands in the fluorescent product within the co-incubated lane (Figure 24, lane 3) indicated at least two TnsB binding sites.

[0155] The systems of this disclosure can be used for a variety of applications, such as nucleic acid editing (e.g., gene editing) or binding to nucleic acid molecules (e.g., sequence-specific binding). Such systems can be used, for example, to improve (e.g., remove or replace) genetic mutations that may cause disease in a target; to inactivate genes to confirm their function within cells; as a diagnostic tool to detect disease-causing genetic elements (e.g., via cleavage of reverse-transcribed viral RNA or amplified DNA sequences encoding disease-causing mutations); as an inactivating enzyme combined with a probe to target and detect specific nucleotide sequences (e.g., sequences encoding antibiotic resistance in bacteria); to inactivate viruses or prevent them from infecting host cells by targeting the viral genome; to add genes or modify metabolic pathways to improve an organism to produce valuable small molecules, macromolecules, or secondary metabolites; to establish gene-driven elements for evolutionary selection; and / or as a biosensor to detect cytotoxicity caused by foreign small molecules and nucleotides.

[0156] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided only as examples. The present invention is not intended to be limited by the specific examples provided herein. Although the present invention is described in relation to the foregoing specification, the descriptions and examples of embodiments herein are not intended to be constrained. Those skilled in the art will be able to conceive of many variations, alterations, and substitutions without departing from the present invention. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific descriptions, configurations, or relative proportions described herein, depending on various conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein may be used in carrying out the present invention. Thus, it is intended that the present invention also encompasses such alternatives, modifications, variations, or equivalents. The following claims define the scope of the present invention, and methods and structures within the scope of these claims and their equivalents are intended to be encompassed thereby.

Claims

1. A manipulated nuclease system, An endonuclease containing a RuvC domain, which is a class II, type V-K Cas effector having at least 90% identity with SEQ ID NO: 1, A manipulated guide ribonucleic acid (RNA) comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, and A modified nuclease system, including [specific component].

2. The manipulated nuclease system according to claim 1, wherein the endonuclease comprises a sequence having at least 90% sequence identity with SEQ ID NO:

1.

3. The manipulated nuclease system according to claim 1 or 2, wherein the endonuclease comprises SEQ ID NO:

1.

4. The manipulated nuclease system according to any one of claims 1 to 3, wherein the manipulated guide RNA comprises a sequence having at least 90% identity with any one of sequence numbers 6, 32-33, 94-95, or 104-105, comprising at least about 46 to 80 consecutive nucleotides.

5. The manipulated nuclease system according to any one of claims 1 to 3, wherein the manipulated guide RNA comprises a sequence having at least 46 to 80 consecutive nucleotides having at least 90% identity with either SEQ ID NO: 5 or 6.

6. The manipulated nuclease system according to any one of claims 1 to 3, wherein the manipulated guide RNA comprises a sequence having at least 90% sequence identity with respect to any one of the undegenerate nucleotides of sequence numbers 106 to 108.

7. The manipulated nuclease system according to any one of claims 1 to 6, wherein the manipulated guide RNA comprises a sequence having at least 90% sequence identity with respect to one of the undegenerate nucleotides of sequence number 5, 45-63, 68-75, or 96-103.

8. The manipulated nuclease system according to any one of claims 1 to 7, wherein the endonuclease is configured to bind to a protospacer adjacent motif (PAM) sequence containing sequence number 31.

9. A double-stranded nucleic acid containing a cargo nucleotide sequence configured to interact with a Tn7 type transposase complex, A Cas effector complex comprising a class II, type V Cas effector and an engineered guide polynucleotide encoding an engineered guide RNA configured to hybridize to a target nucleic acid site, The Tn7 type transposase complex configured to bind to the Cas effector complex, comprising a Tn7 type transposase complex and a Tn7 type transposase complex comprising a TnsB subunit. A system for rearranging the cargo nucleotide sequence to the target nucleic acid site, comprising: A system comprising a class II, type V Cas effector containing a polypeptide having at least 90% identity with sequence number 1.

10. The system according to claim 9, wherein the Class II, Type V Cas effector comprises a polypeptide having at least 90% identity with SEQ ID NO:

1.

11. The system according to claim 9 or 10, wherein the Class II, Type V Cas effector comprises a polypeptide containing Sequence ID No.

1.

12. The system according to any one of claims 9 to 11, wherein the TnsB subunit comprises a polypeptide having a sequence having at least 90% sequence identity with any one of sequence numbers 2, 13, 17, or 65.

13. The system according to any one of claims 9 to 12, wherein the TnsB subunit comprises a polypeptide having a sequence having at least 90% sequence identity with respect to sequence number 2.

14. The system according to any one of claims 9 to 13, wherein the Tn7 type transposase complex comprises a polypeptide having a sequence having at least 90% sequence identity with any one of SEQ ID NOs: 3-4, 14-15, 18-19, or 66-67.

15. The system according to any one of claims 9 to 14, wherein the Tn7 type transposase complex comprises a polypeptide having at least 90% sequence identity with SEQ ID NO: 3 or 4.

16. The system according to any one of claims 9 to 15, wherein the manipulated guide polynucleotide comprises a sequence having at least 90% sequence identity with respect to one of the non-degenerate nucleotides, SEQ ID NOs. 5, 45-63, 68-75, 96-103, 106, 107, or 108.

17. The system according to any one of claims 9 to 16, wherein the manipulated guide polynucleotide comprises a sequence having at least 90% sequence identity with any one of SEQ ID NOs. 5-6, 32-33, 94-95, or 104-105, comprising at least about 46-80 consecutive nucleotides.

18. The system according to any one of claims 9 to 17, wherein the manipulated guide polynucleotide comprises a sequence having at least 90% sequence identity with respect to either SEQ ID NO: 5 or 6, and comprising at least about 46 to 80 consecutive nucleotides.

19. The system according to any one of claims 9 to 18, wherein the cargo nucleotide sequence is adjacent to the left transposase recognition sequence and the right transposase recognition sequence.

20. The system according to claim 19, wherein the transposase recognition sequence on the left side includes a sequence having at least 90% identity with any one of sequence numbers 9, 11, 36-38, 76, or 78.

21. The system according to claim 19 or 20, wherein the transposase recognition sequence on the right side includes a sequence having at least 90% identity with any one of sequence numbers 8, 10, 39-44, 77, 79, or 93.

22. A method for modifying a target nucleotide sequence, comprising the step of contacting the target nucleotide sequence in vitro with the system described in any one of claims 9 to 21.

23. The method according to claim 22, wherein the target nucleotide sequence is located within a cell.

24. One or more nucleic acids encoding the system described in any one of claims 9 to 21.