Systems and methods for rearranging cargo nucleotide sequences
The system employing Class 2 V-type Cas effectors and Tn7-type transposase complexes addresses the challenge of precise cargo nucleotide sequence transposition in nucleic acid sites, improving DNA manipulation and gene editing efficacy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- METAGENOMI THERAPEUTICS INC
- Filing Date
- 2024-05-10
- Publication Date
- 2026-06-02
AI Technical Summary
Current CRISPR/Cas systems for DNA manipulation and gene editing lack efficient methods for transposing cargo nucleotide sequences into specific target nucleic acid sites with high precision and versatility.
A system utilizing a Cas effector complex, a Tn7 transposase complex, and an engineered guide polynucleotide to facilitate the transposition of cargo nucleotide sequences into target nucleic acid sites, leveraging Class 2 V-type Cas effectors and Tn7-type transposase components for precise integration.
Enables precise and versatile transposition of cargo nucleotide sequences into target nucleic acid sites, enhancing the efficiency and accuracy of DNA manipulation and gene editing processes.
Smart Images

Figure 2026517800000001_ABST
Abstract
Description
[Technical Field]
[0001] cross reference This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 501,227, filed on 10 May 2023, and U.S. Provisional Patent Application No. 63 / 514,768, filed on 20 July 2023, each of which is incorporated herein by reference in its entirety.
[0002] Sequence List This application includes a sequence listing submitted electronically in XML format, which is incorporated herein by reference in its entirety. The XML copy, created on 3 May 2024, is named MTG-026WO_SL.xml and is 1,536,000 bytes in size. [Background technology]
[0003] Cas enzymes, along with their associated clustered and regularly arranged short palindromic repeat (CRISPR) guided ribonucleic acid (RNA), appear to be a widespread component of the prokaryotic immune system (approximately 45% of bacteria and 84% of archaea), helping to protect such microorganisms from non-self nucleic acids such as infectious viruses and plasmids through CRISPR-RNA-induced nucleic acid cleavage. While the deoxyribonucleic acid (DNA) elements encoding CRISPR RNA elements can be relatively conserved in structure and length, their CRISPR-associated (Cas) proteins are highly diverse and contain a wide variety of nucleic acid interaction domains. Although CRISPR DNA elements were observed as early as 1987, the programmable endonuclease cleavage capability of the CRISPR / Cas complex has only been recognized relatively recently, leading to the use of recombinant CRISPR / Cas systems in a wide range of DNA manipulation and gene editing applications. [Overview of the Initiative]
[0004] In some embodiments, the Disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site, comprising: a first double-stranded nucleic acid comprising a cargo nucleotide sequence configured to interact with a Tn7 transposase complex; a Cas effector complex comprising a class 2 V-type Cas effector and an engineered guide polynucleotide configured to hybridize to the target nucleotide sequence; and a Tn7 transposase complex configured to bind to the Cas effector complex, wherein the Tn7 transposase complex comprises a TnsB subunit. In some embodiments, the cargo nucleotide sequence is adjacent to a left-side transposase recognition sequence and a right-side transposase recognition sequence. In some embodiments, the system further comprises a second double-stranded nucleic acid comprising the target nucleic acid site. In some embodiments, the system further comprises a PAM sequence adjacent to the target nucleic acid site and compatible with the Cas effector complex. In some embodiments, the PAM sequence is located at 3' of the target nucleic acid site.
[0005] In some embodiments, the present disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, comprising: a Cas effector complex comprising a class 2 V-type Cas effector, a microprokaryotic ribosomal protein subunit S15, and an engineered guide polynucleotide that hybridizes to a target nucleic acid site; a Tn7-type transpose complex that binds to the Cas effector complex, comprising an auxiliary protein containing TnsB, TnsC, and TniQ components, and a sequence having at least 70% sequence homology with any one of SEQ ID NOs. 228-230 and 235-249; and a double-stranded nucleic acid that interacts with the Tn7-type transposase complex and contains a cargo nucleotide sequence.
[0006] In some embodiments, the Cas effector complex is non-covalently bound to the Tn7 transposase complex. In some embodiments, the Cas effector complex is covalently bound to the Tn7 transposase complex. In some embodiments, the Cas effector complex is fused to the Tn7 transposase complex.
[0007] In some embodiments, the cargo nucleotide sequence is adjacent to a left-side transposase recognition sequence and a right-side transposase recognition sequence recognized by the Tn7 type transposase complex. In some embodiments, the left-side transposase sequence includes a sequence having at least 80% identity with any one of sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the right-side transposase sequence includes a sequence having at least 80% identity with any one of sequence numbers 8, 10, 39-44, 77, 79, and 93.
[0008] In some embodiments, the target nucleic acid includes a PAM sequence that fits with the Cas effector complex. In some embodiments, the PAM sequence includes sequence number 31. In some embodiments, the PAM sequence is located about 50 to 70 base pairs from the target nucleic acid site. In some embodiments, the PAM sequence is located at 3' of the target nucleic acid site. In some embodiments, the PAM sequence is located at 5' of the target nucleic acid site.
[0009] In some embodiments, the Class 2 V-type Cas effector is a Cas12k effector. In some embodiments, the Class 2 V-type Cas effector includes a polypeptide having at least 80% identity with any one of sequence numbers 1, 12, 16, 20-30, 64, 80-85, and 220.
[0010] In some embodiments, a Class 2 V-type Cas effector comprises a polypeptide having at least 90% identity with any one of sequence numbers 1, 12, 16, 20-30, 64, 80-85, and 220.
[0011] In some embodiments, the TnsB component includes a polypeptide having a sequence that is at least 80% identical to one of sequence numbers 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide having a sequence that is at least 90% identical to one of sequence numbers 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide having a sequence that is at least 90% identical to one of sequence numbers 2, 13, 17, and 65.
[0012] In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 80% identity with any one of sequence numbers 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 90% identity with any one of sequence numbers 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 90% identity with any one of sequence numbers 3-4, 14-15, 18-19, 66-67, and 109-111.
[0013] In some embodiments, the manipulated guide polynucleotide includes a sequence comprising at least 46 to 80 consecutive nucleotides having at least 80% identity with any one of SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 80% sequence homology with any one of SEQ ID NOs. 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944.
[0014] In some embodiments, the microprokaryotic ribosomal protein subunit S15 includes a sequence having at least 80% sequence homology with any one of sequence numbers 187-189. In some embodiments, the microprokaryotic ribosomal protein subunit S15 is encoded by a sequence having at least 80% sequence homology with any one of sequence numbers 181-183.
[0015] In some embodiments, the class 2 V-type Cas effector and Tn7-type transposase complex is encoded by a polynucleotide sequence of less than approximately 10 kilobases.
[0016] In some embodiments, the auxiliary protein is ClpX, which contains a sequence having at least about 80% sequence homology with any one of sequence numbers 235-249.
[0017] In some embodiments, the present disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, comprising: a Cas effector complex comprising a class 2 V-type Cas effector, a microprokaryotic ribosomal protein subunit S15, and an engineered guide polynucleotide that hybridizes to the target nucleic acid site; a Tn7-type transpotase complex that binds to the Cas effector complex and comprises a functional domain (FD)-TniQ fusion and an auxiliary protein; and a double-stranded nucleic acid that interacts with the Tn7-type transposase complex and comprises a cargo nucleotide sequence.
[0018] In some embodiments, the Cas effector complex is non-covalently bound to the Tn7 transposase complex. In some embodiments, the Cas effector complex is covalently bound to the Tn7 transposase complex. In some embodiments, the Cas effector complex is fused to the Tn7 transposase complex.
[0019] In some embodiments, the present disclosure provides a functional domain (FD) comprising a sequence having at least 70% identity with any one of sequence numbers 257-307 and 1138-1242.
[0020] In some embodiments, the cargo nucleotide sequence is adjacent to a left-side transposase recognition sequence and a right-side transposase recognition sequence recognized by the Tn7 type transposase complex. In some embodiments, the left-side transposase sequence includes a sequence having at least 80% identity with any one of sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the right-side transposase sequence includes a sequence having at least 80% identity with any one of sequence numbers 8, 10, 39-44, 77, 79, and 93.
[0021] In some embodiments, the target nucleic acid comprises a PAM sequence that is compatible with the Cas effector complex. In some embodiments, the PAM sequence comprises SEQ ID NO: 31. In some embodiments, the PAM sequence is located about 50 to about 70 base pairs from the target nucleic acid site. In some embodiments, the PAM sequence is located 3' of the target nucleic acid site. In some embodiments, the PAM sequence is located 5' of the target nucleic acid site.
[0022] In some embodiments, the Class 2 Type V Cas effector is a Cas12k effector. In some embodiments, the Class 2 Type V Cas effector comprises a polypeptide comprising a sequence having at least 90% identity with any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the Class 2 Type V Cas effector comprises a polypeptide comprising a sequence of any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220.
[0023] In some embodiments, the engineered guide polynucleotide comprises a sequence comprising at least about 46 to 80 contiguous nucleotides having at least 80% identity with any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the engineered guide polynucleotide comprises a sequence having at least 80% sequence homology with any one of SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944.
[0024] In some embodiments, the microprokaryotic ribosomal protein subunit S15 comprises a sequence having at least 80% sequence homology with any one of sequence numbers 187-189. In some embodiments, the microprokaryotic ribosomal protein subunit S15 is encoded by a sequence having at least 80% sequence homology with any one of sequence numbers 181-183. In some embodiments, the class 2 V-type Cas effector and Tn7-type transposase complex is encoded by a polynucleotide sequence containing less than approximately 10 kilobases.
[0025] In some embodiments, the auxiliary protein comprises a sequence having at least 70% sequence homology with one of SEQ ID NOs. 228-230 and 235-249. In some embodiments, the auxiliary protein is ClpX comprising a sequence having at least about 80% sequence homology with one of SEQ ID NOs. 235-249.
[0026] In another aspect, the present disclosure provides a system for translocating a cargo nucleotide sequence within a target nucleic acid site in a target nucleic acid, the system comprising a Class 2 type V Cas effector, and an engineered guide polynucleotide that hybridizes to the target nucleic acid site, wherein the Cas effector complex comprises a polypeptide comprising a sequence having at least 80% sequence homology with any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220; a Tn7-type transposase complex that binds to the Cas effector complex and comprises TnsB, TnsC, and TniQ components, wherein the TnsB, TnsC, or TniQ component comprises a sequence having at least 80% sequence homology with any one of SEQ ID NOs: 2-4, 13-15, 17-19, 65-67, and 109-111; a Tn7-type transposase complex comprising an auxiliary protein comprising a sequence having at least 80% sequence homology with any one of SEQ ID NOs: 228-230 and 235-249; and a double-stranded nucleic acid that interacts with the Tn7-type transposase complex and comprises a cargo nucleotide sequence.
[0027] In some embodiments, the Cas effector complex non-covalently binds to the Tn7-type transposase complex. In some embodiments, the Cas effector complex is covalently bound to the Tn7-type transposase complex. In some embodiments, the Cas effector complex is fused to the Tn7-type transposase complex.
[0028] In some embodiments, the cargo nucleotide sequence is adjacent to a left transposase recognition sequence and a right transposase recognition sequence recognized by the Tn7-type transposase complex. In some embodiments, the left transposase recognition sequence comprises a sequence having at least 80% identity with any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some embodiments, the right transposase sequence comprises a sequence having at least 80% identity with any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93.
[0029] In some embodiments, the target nucleic acid includes a PAM sequence that fits with the Cas effector complex. In some embodiments, the PAM sequence includes sequence number 31. In some embodiments, the PAM sequence is located about 50 to 70 base pairs from the target nucleic acid site. In some embodiments, the PAM sequence is located at 3' of the target nucleic acid site. In some embodiments, the PAM sequence is located at 5' of the target nucleic acid site.
[0030] In some embodiments, the Class 2 V-type Cas effector is a Cas12k effector. In some embodiments, the Class 2 V-type Cas effector includes a polypeptide comprising a sequence having at least 90% identity with any one of sequence numbers 1, 12, 16, 20-30, 64, 80-85, and 220.
[0031] In some embodiments, a Class 2 V-type Cas effector comprises a polypeptide containing one of the sequences 1, 12, 16, 20-30, 64, 80-85, and 220.
[0032] In some embodiments, the TnsB, TnsC, or TniQ component includes a sequence having at least 90% sequence homology with any one of sequence numbers 2-4, 13-15, 17-19, 65-67, and 109-111.
[0033] In some embodiments, the manipulated guide polynucleotide includes a sequence comprising at least 46 to 80 consecutive nucleotides having at least 80% identity with any one of SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence having at least 80% sequence homology with any one of SEQ ID NOs. 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944.
[0034] In some embodiments, the microprokaryotic ribosomal protein subunit S15 includes a sequence having at least 80% sequence homology with any one of sequence numbers 187-189. In some embodiments, the microprokaryotic ribosomal protein subunit S15 is encoded by a sequence having at least 80% sequence homology with any one of sequence numbers 181-183.
[0035] In some embodiments, the class 2 V-type Cas effector and Tn7-type transposase complex is encoded by a polynucleotide sequence containing less than approximately 10 kilobases.
[0036] In some embodiments, the auxiliary protein is ClpX, which contains a sequence having at least 90% sequence homology with any one of sequence numbers 235-249.
[0037] In some embodiments, the present disclosure relates to a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, comprising: a Cas effector complex comprising: i) a class 2 V-type Cas effector having at least 80% sequence homology with any one of sequence numbers 1, 81, 82, 83, and 85; and ii) an engineered guide polynucleotide having at least 80% identity with any one of sequence numbers 5, 6, 45-63, 68-75, 96-103, 123-140, and 754-944; and a Tn7-type transposase complex bound to the Cas effector complex, comprising TnsB, TnsC, and TniQ components, and Tn7-type transposase complex, wherein the TnsB, TnsC, or TniQ component is of sequence numbers 2-4. The present invention provides a system comprising: a Tn7 transposase complex comprising a sequence having at least 80% sequence homology with any one of the above, and an auxiliary protein comprising a sequence having at least 80% sequence homology with any one of sequence numbers 228-230 and 235-249; and a double-stranded nucleic acid comprising, in the order from 5' to 3', i) a left transposase recognition sequence comprising a sequence having at least 80% sequence homology with any one of sequence numbers 9, 11, 36, 37, and 38; ii) a cargo nucleotide sequence; and ii) a right transposase sequence comprising a sequence having at least 80% identity with any one of sequence numbers 8, 39-44, and 93.
[0038] In some embodiments, the present disclosure relates to a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, comprising: a Cas effector complex that hybridizes to the target nucleic acid site, comprising: i) a class 2 V-type Cas effector comprising a sequence having at least 80% sequence homology with SEQ ID NO: 12, and iii) an engineered guide polynucleotide having at least 80% identity with any one of SEQ ID NOs: 32, 102, 104, and 107; and a Tn7-type transposase complex that binds to the Cas effector complex and comprises TnsB, TnsC, and TniQ components, wherein TnsB, TnsC, or Tni The present invention provides a system comprising: a Tn7 transposase complex in which the Q component contains a sequence having at least 80% sequence homology with any one of sequence numbers 13 to 15, and the auxiliary protein contains a sequence having at least 80% sequence homology with any one of sequence numbers 228 to 230 and 235 to 249; and a double-stranded nucleic acid that interacts with the Tn7 transposase complex and contains, in the order from 5' to 3', i) a left-side transposase sequence containing a sequence having at least 80% sequence homology with sequence number 76, ii) the cargo nucleotide sequence, and iii) a right-side transposase sequence containing a sequence having at least 80% identity with sequence number 77.
[0039] In some embodiments, the present disclosure relates to a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, comprising: a Cas effector complex that hybridizes to the target nucleic acid site, comprising: i) a class 2 V-type Cas effector comprising a sequence having at least 80% sequence homology with SEQ ID NO: 16, and ii) an engineered guide polynucleotide having at least 80% identity with any one of SEQ ID NOs: 33, 103, 105, and 108; and a Tn7-type transposase complex that binds to the Cas effector complex and comprises TnsB, TnsC, and TniQ components, wherein the TnsB, TnsC, or TniQ component is sequence The present invention provides a system comprising: a Tn7 transposase complex comprising a sequence having at least 80% sequence homology with any one of sequence numbers 17-19, and an auxiliary protein comprising a sequence having at least 80% sequence homology with any one of sequence numbers 228-230 and 235-249; and a double-stranded nucleic acid that interacts with the Tn7 transposase complex and comprises, in the order from 5' to 3', i) a left transposase recognition sequence comprising a sequence having at least 80% sequence homology with sequence number 78, ii) a cargo nucleotide sequence, ii) another cargo nucleotide sequence, and iii) a right transposase sequence comprising a sequence having at least 80% identity with sequence number 79.
[0040] In some embodiments, the system further comprises a PAM sequence compatible with the Cas effector complex. In some embodiments, the PAM sequence includes sequence number 31.
[0041] In some embodiments, the PAM sequence is located approximately 50 to 70 base pairs from the target nucleic acid site. In some embodiments, the PAM sequence is located at 3' of the target nucleic acid site. In some embodiments, the PAM sequence is located at 5' of the target nucleic acid site.
[0042] In some embodiments, the Cas effector complex further comprises a microprokaryotic ribosomal protein subunit S15. In some embodiments, the microprokaryotic ribosomal protein subunit S15 comprises a sequence having at least 80% sequence homology with any one of sequence numbers 187-189. In some embodiments, the auxiliary protein is ClpX comprising a sequence having at least about 80% sequence homology with any one of sequence numbers 235-249.
[0043] In some embodiments, the present disclosure provides an engineered nuclease system comprising: an endonuclease comprising a RuvC domain, which is a class 2 VK-type Cas effector derived from an uncultured microorganism and having at least 80% identity with any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220; and an engineered guide polynucleotide comprising a sparser sequence that forms a complex with the endonuclease and hybridizes to a target nucleic acid sequence, wherein the engineered guide polynucleotide comprises a sequence having at least about 80% identity with any one of SEQ ID NOs: 754-944.
[0044] In some embodiments, the Disclosure provides a method for transposing a cargo nucleotide sequence into a target nucleic acid site, comprising introducing one of the systems of the Disclosure into a cell. In some embodiments, the Disclosure provides a cell comprising one of the systems of the Disclosure. In some embodiments, the cell is a eukaryotic cell.
[0045] In some embodiments, the cells are mammalian cells. In some embodiments, the cells are immortalized cells. In some embodiments, the cells are insect cells. In some embodiments, the cells are yeast cells. In some embodiments, the cells are plant cells. In some embodiments, the cells are fungal cells. In some embodiments, the cells are prokaryotic cells. In some embodiments, the cells are A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cells, HT1080, HepG2, Huh7, K562, primary cells, or derivatives thereof. In some embodiments, the cells are engineered cells. In some embodiments, the cells are stable cells.
[0046] Further aspects and advantages of the present disclosure will be readily apparent to those skilled in the art from the following detailed description, which shows and describes only exemplary embodiments of the present disclosure. As will be recognized, other different embodiments of the present disclosure are possible, and some of their details can be modified in various obvious ways without departing from the present disclosure. Accordingly, the drawings and description should be considered illustrative in nature, not restrictive.
[0047] Novel features of this disclosure are specifically described in the appended claims. A better understanding of the features and advantages of this disclosure will be obtained by referring to the following detailed description, which describes illustrative embodiments in which the principles of this disclosure are utilized, and to the appended drawings (also referred to herein as "Figure" and "FIG."). [Brief explanation of the drawing]
[0048] [Figure 1] This paper describes exemplary mechanisms of different classes and types of CRISPR / Cas loci. [Figure 2]This describes the structure of a native class 2 type II crRNA / tracrRNA pair, as shown in Cas9, compared to a hybrid sgRNA to which crRNA and tracrRNA are joined. [Figure 3] Two paths found in Tn7 and Tn7-like elements are illustrated. [Figure 4A] The genomic context of the V-type Tn7 CAST of the MG64 family is illustrated. Figure 4A depicts the MG64-1 CAST system as containing a CRISPR array (CRISPR repeats), a V-type nuclease, and three predicted transposase protein sequences. The tracrRNA was predicted within the intergenetic region between the CAST effector array and the CRISPR array. Below: Multiple sequence alignment of the catalytic domain of transposase TnsB. Catalytic residues are indicated by boxes. [Figure 4B] The genomic context of the V-type Tn7 CAST of the MG64 family is illustrated. Figure 4B depicts the predictions for the two transposon ends for the MG64-1 CAST system. [Figure 5] This section depicts the predicted structures of the corresponding sgRNAs of the CAST system described herein. Panel A (left) of Figure 5 shows the predicted MG64-1 tracrRNA and crRNA double-strand complex in a repeat-anti-repeat stem. The loop was cleaved, and a tetraloop of GAAA was added to the stem-loop structure to produce the designed sgRNA shown in Panel B (right) of Figure 5. [Figure 6] The results of a rearrangement reaction targeting a plasmid library consisting of NNNNNNNN at the 5' of the target spacer sequence are depicted. Reaction #1 shows the presence of the target library, #2 shows the presence of donor fragments in both rearrangement reactions, and #3-5 show sg-specific PCR bands corresponding to the appropriate rearrangement reactions. [Figure 7A]The results of Sanger sequencing are depicted. Figure 7A shows the Sanger sequencing of the donor-target junction on the left end (LE) of a transposon in an LE rearrangement reaction closer to the PAM. The expected sequence is at the top of the panel, accompanied by the predicted rearrangement event 61 bp away from the PAM. The upper chromatogram is the sequencing result starting from within the donor fragment. A clear signal is seen at the right end up to the donor / target junction (dotted line). This indicates a mixture of rearrangement products. The lower chromatogram is the sequencing from the target to the donor / target junction. The signal from the left is a clear signal up to the junction. [Figure 7B] The results of Sanger sequencing are depicted. Figure 7B shows the Sanger sequencing of the donor-target junction on the right end (RE) of the transposon in the LE product closer to the PAM. The expected sequence is at the top of the panel, accompanied by a predicted dislocation event 61 bp away from the PAM. The chromatogram above shows the sequencing result starting from within the donor fragment. A clear signal is seen at the left end up to the donor / target junction (dotted line). [Figure 7C] The results of Sanger sequencing are depicted. Figure 7C is a close-up view of the PAM library. [Figure 7D] The results of Sanger sequencing are depicted. Figure 7D shows a SeqLogo analysis of NGS for LE events that are closer to PAM, showing a very strong preference for NGTN in the PAM motif. [Figure 8] This diagram depicts the phylogenetic gene lineage of Cas12k effector sequences. This lineage was inferred from multiple sequence alignments of 64 Cas12k sequences recovered here (orange and black branches) and 229 reference Cas12k sequences from publicly available databases (gray branches). Orange branches indicate Cas12k effectors whose association with CAST transposon components has been confirmed. [Figure 9]The MG64 family CRISPR repeat alignment is shown. The Cas12k CAST CRISPR repeat contains the conserved motif 5'-GNNGGNNTGAAAG-3'. In MG64-1, the short repeat-repeat repression (RAR) within the CRISPR repeat motif aligns with the tracrRNA. The MG64 RAR motif appears to define the start and end of the tracrRNA (5' end: RAR1(TTTC), 3' end: RAR2(CCNNC)). [Figure 10A-1] The secondary structure predicted from the folding of CRISPR repeats + tracrRNA on the MG64 system is illustrated. [Figure 10A-2] The secondary structure predicted from the folding of CRISPR repeats + tracrRNA on the MG64 system is illustrated. [Figure 10B-1] The secondary structure predicted from the folding of CRISPR repeats + tracrRNA on the MG64 system is illustrated. [Figure 10B-2] The secondary structure predicted from the folding of CRISPR repeats + tracrRNA on the MG64 system is illustrated. [Figure 11A] The MG64-3 CRISPR locus is illustrated. TracrRNA is encoded upstream of the CRISPR array, while the transposon end is encoded downstream (inner black box). Sequences corresponding to partial 3' CRISPR repeats and partial spacers are encoded within the transposon (outer box). Self-matching spacers are encoded outside the transposon end. [Figure 11B] The tracrRNA sequence alignments for various CASTs provided herein are illustrated. The tracrRNA sequence alignments indicate conserved regions. In particular, the sequence "TGCTTTC" (upper box) at positions 92–98 may be important for sgRNA tertiary structure and for discontinuous repeat-repeat repression pairing with crRNA. The hairpin "CYCC(n6)GGRG" (lower box) at positions 265–278 may be important for function, such as positioning downstream sequences for crRNA pairing. [Figure 12A] This diagram depicts the predicted structure of MG64-1 sgRNA. [Figure 12B] The predicted structure of MG64-3 sgRNA is shown in the diagram. [Figure 12C] Describe the predicted structure of MG64-5 sgRNA. [Figure 13A] This section depicts PCR data demonstrating the activity of MG64-1 with sgRNA v2-1. Effector proteins, along with their TnsB, TnsC, and TniQ proteins, were expressed in an in vitro transcription / translation system using the protocol described for in vitro targeted integrase activity. Post-translation, target DNA, cargo DNA, and sgRNA were added to the reaction buffer. Integration was assayed by PCR across the target / donor junction. Figure 13A illustrates a schematic diagram showing the potential orientation of the integrated donor DNA. PCR reactions 3, 4, 5, and 6 represent the respective integration ligation products depending on the orientation of the donor at the target site. [Figure 13B] The PCR data illustrating the activity of MG64-1 with sgRNA v2-1 is depicted. Using the protocol described for in vitro targeted integrase activity, the effector protein and its TnsB, TnsC, and TniQ proteins were expressed in an in vitro transcription / translation system. Post-translation, target DNA, cargo DNA, and sgRNA were added to the reaction buffer. Integration was assayed by PCR across the target / donor junction. Figure 13B illustrates the gel image of PCR 4 (detecting RE junctions to the donor) of transposition, showing lane 1) apo (no sgRNA), lane 2) with sgRNA 1, and lane 3) with sgRNA v2-1. [Figure 13C]The PCR data illustrating the activity of MG64-1 with sgRNA v2-1 is depicted. Using the protocol described for in vitro targeted integrase activity, the effector protein and its TnsB, TnsC, and TniQ proteins were expressed in an in vitro transcription / translation system. Post-translation, target DNA, cargo DNA, and sgRNA were added to the reaction buffer. Integration was assayed by PCR across the target / donor junction. Figure 13C illustrates the gel image of PCR 5 (detecting the LE junction to the donor) of transposition, showing lane 1) apo (no sgRNA), lane 2) with sgRNA 1, and lane 3) with sgRNA v2-1. [Figure 14] PCR reaction 5 (LE proximal to PAM, upper half of the plot) and PCR reaction 4 (RE distal to PAM, lower half of the plot) of MG64-1 plotted against the sequence and distance from PAM are shown. Analysis of the integration window shows that 95% of integrations occurring at the spacer PAM site are within a 10 bp window, 58–68 nucleotides away from PAM. The difference in integration distance between distal and proximal frequencies reflects overlap of integration sites, i.e., 3–5 base pairs of overlap as a result of alternating nuclease activity of the transposases during integration. [Figure 15] The results of colony PCR screening for rearrangement efficiency are shown in the figure. After incubation, 18 colony-forming units (CFUs) were visible on the plate, with 8 on plate A (without IPTG, labeled as A) and 10 on plate B (with 100 μM IPTG during recovery, labeled as B). All 18 were analyzed by colony PCR, which yielded a product band indicating successful rearrangement (arrow). [Figure 16]The sequencing results of selected colony PCR products are illustrated, confirming that they represent transposition events because they extend to the junction between the LE and PAM at the manipulated target site within the lacZ gene. The minimum LE sequence is shown in blue at the top of the screen (minimum LE), while the target and PAM are shown in gray. Some sequence variation is observed in the PCR products, which is expected considering that insertions can occur at variable distances upstream of the PAM. [Figure 17]64-1 The results of the experiment with manipulated single guides for translocation activity are illustrated. Black boxes indicate lanes not relevant to this experiment. Panel A of Figure 17 illustrates the gel image of PCR 4 for translocation (detecting RE junctions to donors): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = sgRNA v1-1, Lane 4 = sgRNA v1-2, Lane 5 = sgRNA v1-3. Panel B of Figure 17 illustrates the gel image of PCR 5 for translocation (detecting LE junctions to donors): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = sgRNA v1-1, Lane 4 = sgRNA v1-2, Lane 5 = sgRNA v1-3. Panel C of Figure 17 illustrates the gel image of PCR4 for translocation (detecting the RE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = sgRNA v1-4, Lane 4 = sgRNA v1-6, Lane 5 = sgRNA v1-7, Lane 6 = sgRNA v1-8, Lane 7 = sgRNA v1-9. Panel D of Figure 17 illustrates the gel image of PCR5 for translocation (detecting the LE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = sgRNA v1-4, Lane 4 = sgRNA v1-6, Lane 5 = sgRNA v1-7, Lane 6 = sgRNA v1-8, Lane 7 = sgRNA v1-9. Panel E in Figure 17 illustrates the gel image of PCR4 for translocation (detecting the RE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = sgRNA v1-5, Lane 4 = skip, Lane 5 = sgRNA v1-10. Panel F in Figure 17 illustrates the gel image of PCR5 for translocation (detecting the LE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = sgRNA v1-5, Lane 4 = skip, Lane 5 = sgRNA v1-10.Panel G of Figure 17 illustrates the gel image of PCR 4 for translocation (detecting the RE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = sgRNA v1-17, Lane 4 = sgRNA v1-18, Lane 5 = skipped, Lane 6 = sgRNA v1-19, Lane 7 = skipped, Lane 8 = sgRNA v1-20. Panel H of Figure 17 illustrates the gel image of PCR 5 for translocation (detecting the LE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = sgRNA v1-17, Lane 4 = sgRNA v1-18, Lane 5 = skipped, Lane 6 = sgRNA v1-19, Lane 7 = skipped, Lane 8 = sgRNA v1-20. [Figure 18]64-1 Results of the tests of manipulated LE and RE for translocation activity are shown. Black boxes are lanes not relevant to this experiment. Panel A of Figure 18 depicts the gel image of PCR4 for translocation (detecting RE junctions to donors): Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+ sgRNA), Lane 3 = LE 86 bp, Lane 4 = LE 105 bp, Lane 5 = RE 196 bp, Lane 6 = RE 242 bp, Lane 7 = RE internal deletion 50, Lane 8 = RE internal deletion 81. Panel B of Figure 18 depicts the gel image of PCR 5 for translocation (detecting the LE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = LE 86bp, Lane 4 = LE 105bp, Lane 5 = RE 196bp, Lane 6 = RE 242bp, Lane 7 = RE internal deletion 50, Lane 8 = RE internal deletion 81. Panel C of Figure 18 illustrates the gel image of PCR 4 for translocation (detecting the RE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = RE internal deletion 81 and 178bp, Lane 4 = skipped, Lane 5 = RE internal deletion 81 and 196bp, Lane 6 = skipped, Lane 7 = RE internal deletion 81 and 212bp, Lane 8 = skipped. Panel D of Figure 18 shows the gel image of PCR 5 for translocation (detecting the LE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = RE internal deletions 81 and 178 bp, Lane 4 = skipped, Lane 5 = RE internal deletions 81 and 196 bp, Lane 6 = skipped, Lane 7 = RE internal deletions 81 and 212 bp, Lane 8 = skipped. Panel E of Figure 18 shows the gel image of PCR 4 for translocation (detecting the RE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = RE internal deletions 81 and 178 bp + LE 68 bp, Lane 4 = RE internal deletions 81 and 178 bp + LE 86 bp, Lane 5 = skipped, Lane 6 = RE internal deletions 81 and 178 bp + LE 105 bp, Lane 7 = skipped.Panel F of Figure 18 illustrates the gel image of PCR 5 for translocation (detecting the LE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = RE internal deletion 81 and 178 bp + LE 68 bp, Lane 4 = RE internal deletion 81 and 178 bp + LE 86 bp, Lane 5 = skipped, Lane 6 = RE internal deletion 81 and 178 bp + LE 105 bp, Lane 7 = skipped. Panel G of Figure 18 illustrates the gel image of PCR 6 for translocation (detecting the RE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = 0 bp overhang, Lane 4 = 1 bp overhang, Lane 5 = 2 bp overhang, Lane 6 = 3 bp overhang, Lane 7 = 5 bp overhang, Lane 8 = 10 bp overhang. [Figure 19]The results of testing manipulated CAST components, including NLS, for translocation activity are illustrated. Black boxes indicate lanes not relevant to this experiment. Panel A of Figure 19 illustrates the gel image of PCR4 for translocation (detecting RE junctions to donors): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = skipped, Lane 4 = skipped, Lane 5 = skipped, Lane 6 = NLS-TnsB, Lane 7 = skipped, Lane 8 = TnsB-NLS. Panel B of Figure 19 illustrates the gel image of PCR5 for translocation (detecting LE junctions to donors): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = skipped, Lane 4 = skipped, Lane 5 = skipped, Lane 6 = NLS-TnsB, Lane 7 = skipped, Lane 8 = TnsB-NLS. Panel C of Figure 19 illustrates the gel image of PCR4 for translocation (detecting the RE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = skipped, Lane 4 = skipped, Lane 5 = skipped, Lane 6 = NLS-TniQ, Lane 7 = skipped, Lane 8 = TniQ-NLS. Panel D of Figure 19 illustrates the gel image of PCR5 for translocation (detecting the LE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = skipped, Lane 4 = skipped, Lane 5 = skipped, Lane 6 = NLS-TniQ, Lane 7 = skipped, Lane 8 = TniQ-NLS. Panel E in Figure 19 illustrates the gel image of PCR 4 for translocation (detecting the RE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = skipped, Lane 4 = skipped, Lane 5 = NLS-Cas12k, Lane 6 = Cas12k-NLS, Lane 7 = NLS-TnsC, Lane 8 = TnsC-NLS. Panel F in Figure 19 illustrates the gel image of PCR 5 for translocation (detecting the LE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = skipped, Lane 4 = skipped, Lane 5 = NLS-Cas12k, Lane 6 = Cas12k-NLS, Lane 7 = NLS-TnsC, Lane 8 = TnsC-NLS.Panel G in Figure 19 illustrates the gel image of PCR4 for translocation (detecting the RE junction to the donor): Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+ sgRNA), Lane 3 = NLS-HA-TnsC, Lane 4 = NLS-TnsC-FLAG, Lane 5 = NLS-TnsC-HA, Lane 6 = NLS-TnsC-Myc, Lane 7 = NLS-FLAG-TnsC, Lane 8 = NLS-Myc-TnsC. Panel H in Figure 19 illustrates the gel image of PCR 5 for translocation (detecting the LE junction to the donor): Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = NLS-HA-TnsC, Lane 4 = NLS-TnsC-FLAG, Lane 5 = NLS-TnsC-HA, Lane 6 = NLS-TnsC-Myc, Lane 7 = NLS-FLAG-TnsC, Lane 8 = NLS-Myc-TnsC. Panel I in Figure 19 illustrates the gel image of PCR 4 for translocation (detecting the RE junction to the donor): Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = Cas 2x NLS Apo (no sgRNA), Lane 4 = Cas 2x NLS Holo (+sgRNA). Panel J in Figure 19 illustrates the gel image of PCR5 for translocation (detecting the LE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = Cas 2x NLS apo (no sgRNA), Lane 4 = Cas 2x NLS holo (+sgRNA). [Figure 20]The manipulated CAST-NLS, acting as a single suite, is illustrated. All lanes, unless otherwise noted, contain Cas12k-NLS, and NLS-TniQ, TnsB, TnsC, and sgRNA. Panel A of Figure 20 illustrates the gel image of PCR4 for translocation (detecting RE junctions to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = NLS-TnsB, Lane 4 = TnsB-NLS, Lane 5 = NLS-TnsB and NLS-TnsC, Lane 6 = TnsB-NLS and NLS-TnsC. Panel B in Figure 20 shows gel images of PCR5 for translocation (detecting LE junctions to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+ sgRNA), Lane 3 = NLS-TnsB, Lane 4 = TnsB-NLS, Lane 5 = NLS-TnsB and NLS-TnsC, Lane 6 = TnsB-NLS and NLS-TnsC. [Figure 21]The results of the Cas effector and TniQ protein fusion tests on translocation activity are illustrated. Panel A of Figure 21 shows the gel image of PCR 4 for translocation (detecting the RE junction to the donor): Lane 1 = apo with Cas-TniQ fusion (no sgRNA), Lane 2 = holo with Cas-TniQ fusion (+sgRNA), Lane 3 = apo with TniQ-Cas fusion (no sgRNA), Lane 4 = holo with TniQ-Cas fusion (+sgRNA). Panel B of Figure 21 shows the gel image of PCR 5 for translocation (detecting the LE junction to the donor): Lane 1 = apo with Cas-TniQ fusion (no sgRNA), Lane 2 = holo with Cas-TniQ fusion (+sgRNA), Lane 3 = apo with TniQ-Cas fusion (no sgRNA), Lane 4 = holo with TniQ-Cas fusion (+sgRNA). Panel C in Figure 21 illustrates the gel image of PCR4 for translocation (detecting the RE junction to the donor): Lane 1 = apo with TniQ-Cas fusion (no sgRNA), Lane 2 = holo with TniQ-Cas fusion (+sgRNA), Lane 3 = holoCas alone, Lane 4 = apo with TniQ-48 linker-Cas fusion (no sgRNA), Lane 5 = holo with TniQ-48 linker-Cas fusion (+sgRNA), Lane 6 = apo with TniQ-68 linker-Cas fusion (no sgRNA), Lane 7 = holo with TniQ-68 linker-Cas fusion (+sgRNA), Lane 8 = holo with TniQ-72 linker-Cas fusion (+sgRNA). Panel D in Figure 21 illustrates the gel images of PCR5 for translocation (detecting the LE junction to the donor): Lane 1 = Apo with TniQ-Cas fusion (no sgRNA), Lane 2 = Holo with TniQ-Cas fusion (+sgRNA), Lane 3 = HoloCas alone, Lane 4 = Apo with TniQ-48 linker-Cas fusion (no sgRNA), Lane 5 = Holo with TniQ-48 linker-Cas fusion (+sgRNA), Lane 6 = Apo with TniQ-68 linker-Cas fusion (no sgRNA), Lane 7 = Holo with TniQ-68 linker-Cas fusion (+sgRNA), Lane 8 = Holo with TniQ-72 linker-Cas fusion (+sgRNA).Panel E in Figure 21 illustrates the gel image of PCR4 for translocation (detecting the RE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = apo with NLS-TniQ-Cas-NLS fusion (no sgRNA), Lane 4 = holo with NLS-TniQ-Cas-NLS fusion (+sgRNA), Lane 5 = apo with NLS-TniQ-77 linker-Cas-NLS fusion (no sgRNA), Lane 6 = holo with NLS-TniQ-77 linker-Cas-NLS fusion (+sgRNA). Panel F in Figure 21 illustrates the gel image of PCR5 for translocation (detecting the LE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = apo with NLS-TniQ-Cas-NLS fusion (no sgRNA), Lane 4 = holo with NLS-TniQ-Cas-NLS fusion (+sgRNA), Lane 5 = apo with NLS-TniQ-77 linker-Cas-NLS fusion (no sgRNA), Lane 6 = holo with NLS-TniQ-77 linker-Cas-NLS fusion (+sgRNA). Panel G in Figure 21 illustrates the gel image of PCR4 for translocation (detecting the RE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = NLS-TniQ-Cas-NLS apo (no sgRNA), Lane 4 = NLS-TniQ-Cas-NLS holo (+sgRNA), Lane 5 = Cas-NLS-P2A-NLS-TniQ apo (no sgRNA), Lane 6 = Cas-NLS-P2A-NLS-TniQ holo (+sgRNA). Panel H in Figure 21 illustrates the gel image of PCR5 for translocation (detecting the LE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = NLS-TniQ-Cas-NLS apo (no sgRNA), Lane 4 = NLS-TniQ-Cas-NLS holo (+sgRNA), Lane 5 = Cas-NLS-P2A-NLS-TniQ apo (no sgRNA), Lane 6 = Cas-NLS-P2A-NLS-TniQ holo (+sgRNA). [Figure 22]The expression of TnsB and TnsC in human cells, followed by the results of cell fractionation and in vitro translocation reactions, are illustrated. Panel A of Figure 22 illustrates the gel images of PCR4 for translocation (detecting the RE junction to the donor): Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = Holo with untreated (no TnsB) cytoplasm (+sgRNA), Lane 4 = Holo with untreated nucleoplasm (+sgRNA), Lane 5 = Holo with NLS-TnsB cytoplasm (+sgRNA), Lane 6 = Holo with NLS-TnsB nucleoplasm (+sgRNA), Lane 7 = Holo with TnsB-NLS cytoplasm (+sgRNA), Lane 8 = Holo with TnsB-NLS nucleoplasm (+sgRNA), Lane 9 = Holo with NLS-TniQ cytoplasm (+sgRNA), Lane 10 = Holo with NLS-TniQ nucleoplasm (+sgRNA). Panel B in Figure 22 illustrates the gel image of PCR5 for translocation (detecting the LE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = holo with untreated (no TnsB) cytoplasm (+sgRNA), Lane 4 = holo with untreated nucleoplasm (+sgRNA), Lane 5 = holo with NLS-TnsB cytoplasm (+sgRNA), Lane 6 = holo with NLS-TnsB cell nucleoplasm (+sgRNA), Lane 7 = holo with TnsB-NLS cytoplasm (+sgRNA), Lane 8 = holo with TnsB-NLS cell nucleoplasm (+sgRNA), Lane 9 = holo with NLS-TniQ cytoplasm (+sgRNA), Lane 10 = holo with NLS-TniQ cell nucleoplasm (+sgRNA). Panel C in Figure 22 illustrates the gel image of PCR4 for translocation (detecting the RE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = holo without TnsC (+sgRNA), Lane 4 = holo with untreated (no TnsC) cytoplasm (+sgRNA), Lane 5 = holo with untreated nucleoplasm (+sgRNA), Lane 6 = holo with NLS-HA-TnsC cytoplasm (+sgRNA), Lane 7 = holo with NLS-HA-TnsC cell nucleoplasm (+sgRNA), Lane 8 = holo with TnsC-NLS cytoplasm (+sgRNA), Lane 9 = holo with TnsC-NLS cell nucleoplasm (+sgRNA).Panel D in Figure 22 illustrates the gel image of PCR5 for translocation (detecting the LE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = holo without TnsC (+sgRNA), Lane 4 = holo with untreated (no TnsC) cytoplasm (+sgRNA), Lane 5 = holo with untreated nucleoplasm (+sgRNA), Lane 6 = holo with NLS-HA-TnsC cytoplasm (+sgRNA), Lane 7 = holo with NLS-HA-TnsC cell nucleoplasm (+sgRNA), Lane 8 = holo with TnsC-NLS cytoplasm (+sgRNA), Lane 9 = holo with TnsC-NLS cell nucleoplasm (+sgRNA). Panel E in Figure 22 illustrates the gel image of PCR4 for translocation (detecting the RE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = apo (no sgRNA) NLS-TnsB-IRES-NLS-TnsC cytoplasm, Lane 4 = holo (+sgRNA) NLS-TnsB-IRES-NLS-TnsC cytoplasm, Lane 5 = apo (no sgRNA) NLS-TnsB-IRES-NLS-TnsC nucleoplasm, Lane Lane 6 = Holo(+sgRNA)NLS-TnsB-IRES-NLS-TnsC nucleoplasm, Lane 7 = Apo(no sgRNA)TnsB-NLS-IRES-NLS-TnsC cytoplasm, Lane 8 = Holo(+sgRNA)TnsB-NLS-IRES-NLS-TnsC cytoplasm, Lane 9 = Apo(no sgRNA)TnsB-NLS-IRES-NLS-TnsC nucleoplasm, Lane 10 = Holo(+sgRNA)TnsB-NLS-IRES-NLS-TnsC nucleoplasm.Panel F in Figure 22 illustrates the gel image of PCR5 for translocation (detecting the LE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = apo (no sgRNA) NLS-TnsB-IRES-NLS-TnsC cytoplasm, Lane 4 = holo (+sgRNA) NLS-TnsB-IRES-NLS-TnsC cytoplasm, Lane 5 = apo (no sgRNA) NLS-TnsB-IRES-NLS-TnsC nucleoplasm, Lane Lane 6 = Holo(+sgRNA)NLS-TnsB-IRES-NLS-TnsC nucleoplasm, Lane 7 = Apo(no sgRNA)TnsB-NLS-IRES-NLS-TnsC cytoplasm, Lane 8 = Holo(+sgRNA)TnsB-NLS-IRES-NLS-TnsC cytoplasm, Lane 9 = Apo(no sgRNA)TnsB-NLS-IRES-NLS-TnsC nucleoplasm, Lane 10 = Holo(+sgRNA)TnsB-NLS-IRES-NLS-TnsC nucleoplasm. [Figure 23]The expression of Cas12k and TniQ binding constructs in human cells, followed by the results of in vitro translocation studies, are illustrated. Panel A of Figure 23 illustrates the gel images of PCR5 for translocation (detecting the LE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = Cas-NLS holo (+sgRNA) cytoplasm, Lane 4 = Cas-NLS holo (+sgRNA) nucleoplasm, Lane 5 = Cas-NLS holo (+sgRNA) nucleoplasm + additional sgRNA, Lane 6 = Cas-NLS-P2A-NLS-TniQ holo (+sgRNA) cytoplasm, Lane 7 = Cas-NLS-P2A-NLS-TniQ holo (+sgRNA) cytoplasm, Lane 8 = Cas-NLS-P2A-NLS-TniQ holo (+sgRNA) cytoplasm + additional sgRNA. Panel B of Figure 23 illustrates the gel image of PCR4 for translocation (detecting the RE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = apo (no sgRNA) Cas-NLS-P2A-NLS-TniQ cytoplasm, Lane 4 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ cytoplasm, Lane 5 = apo (no sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm, Lane 6 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm, Lane 7 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm + additional holo-Cas-NLS, Lane 8 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm + NLS-TniQ.Panel C in Figure 23 illustrates the gel image of PCR5 for translocation (detecting the LE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = apo (no sgRNA) Cas-NLS-P2A-NLS-TniQ cytoplasm, Lane 4 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ cytoplasm, Lane 5 = apo (no sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm, Lane 6 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm, Lane 7 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm + additional holo-Cas-NLS, Lane 8 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm + NLS-TniQ. Panel D in Figure 23 illustrates the gel image of PCR4 for translocation (detecting the RE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = apo (no sgRNA) NLS-TniQ-Cas-NLS cytoplasm, Lane 4 = holo (+sgRNA) NLS-TniQ-Cas-NLS cytoplasm, Lane 5 = apo (no sgRNA) NLS-TniQ-Cas-NLS nucleoplasm, Lane 6 = holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm, Lane 7 = holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm + additional holoCas-NLS, Lane 8 = holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm + NLS-TniQ. Panel E in Figure 23 illustrates the gel image of PCR5 for translocation (detecting the LE junction to the donor): Lane 1 = apo (no sgRNA), Lane 2 = holo (+sgRNA), Lane 3 = apo (no sgRNA) NLS-TniQ-Cas-NLS cytoplasm, Lane 4 = holo (+sgRNA) NLS-TniQ-Cas-NLS cytoplasm, Lane 5 = apo (no sgRNA) NLS-TniQ-Cas-NLS nucleoplasm, Lane 6 = holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm, Lane 7 = holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm + additional holoCas-NLS, Lane 8 = holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm + NLS-TniQ.Panel F in Figure 23 shows gel images of PCR4 for transposition (detecting RE junctions to donors): Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = Apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ cytoplasm, Lane 4 = Holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ cytoplasm, Lane 5 = Apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm, Lane 6 = Apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional PURExpress, Lane 7 = Apo (no sgRNA) Cas-N LS-IRES-NLS-TniQ nucleoplasm + additional Cas-NLS, lane 8 = apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + NLS-TniQ, lane 9 = holo(+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm, lane 10 = holo(+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional PURExpress, lane 11 = holo(+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional Cas-NLS, lane 12 = holo(+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + NLS-TniQ.Panel G in Figure 23 illustrates the gel image of PCR5 for translocation (detecting the LE junction to the donor): Lane 1 = Apo (no sgRNA), Lane 2 = Holo (+sgRNA), Lane 3 = Apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ cytoplasm, Lane 4 = Holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ cytoplasm, Lane 5 = Apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm, Lane 6 = Apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional PURExpress, Lane 7 = Apo (no sgRNA) Cas- NLS-IRES-NLS-TniQ nucleoplasm + additional Cas-NLS, lane 8 = apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + NLS-TniQ, lane 9 = holo(+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm, lane 10 = holo(+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional PURExpress, lane 11 = holo(+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional Cas-NLS, lane 12 = holo(+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + NLS-TniQ. [Figure 24] 64-1 The results of electrophoretic mobility shift assays (EMSA) of TnsB and its LE DNA sequence are illustrated. EMSA results confirm binding and TnsB recognition. TnsB protein was expressed in an in vitro transcription / translation system, incubated with FAM-labeled DNA containing the LE sequence, and then separated on a native 5% TBE gel. Binding is observed as an upward shift in the labeled band. Multiple TnsB binding sites lead to multiple shifts in EMSA. Lane 1: FAM-labeled DNA only. Lane 2: FAM DNA + in vitro transcription / translation system (no TnsB protein). Lane 3: FAM DNA + TnsB. [Figure 25A]This depicts Cas12k effector diversity. Figure 25A depicts the Cas12k CAST genomic context. Transposons are characterized by terminal inversion repeats (TIRs, light orange bars), Tn7-like transposon genes (colored arrows), dead effector Cas12k (orange arrows), tracrRNA (pink half-arrows), and CRISPR arrays. "TAAA" target site duplication (TDS) was observed adjacent to the TIR. Center panel: Inset of the MG64-1 non-coding region showing tracrRNA, pseudo-repeats and self-targeting spacers, CRISPR arrays, and transposon left-end TIR. Bottom panel: Multiple alignment of pseudo-repeats and self-targeting spacers in a group of CAST homologs. [Figure 25B] Figure 25B depicts the diversity of Cas12k effectors. It shows an unrooted phylogenetic tree of Cas12k effectors. Cas12k effectors recovered in this study are shown as orange (confirmed transposons in the genome) and black branches, while reference Cas12k sequences are shown in gray. The reference sequences ShCas12k and AcCas12k are indicated by red arrows. [Figure 26A] The multi-array alignment at the right end of the CAST is depicted. The transposon terminal inversion motif "TGTNNA" is highlighted with a box. [Figure 26B] The multi-array alignment at the left end of the CAST is depicted. The transposon terminal inversion motif "TGTNNA" is highlighted with a box. [Figure 27] The alignment of the Cas12k CAST tracrRNA sequence is depicted, showing regions of sequence and structural conservation. In particular, the sequence "TGCTTTC" at positions 88–92 may be important for the sgRNA tertiary structure and for discontinuous repeat-antirepeat pairing with crRNA. The hairpin "CYCC(n6)GGRG" at positions 279–294 may be important for function, potentially positioning downstream sequences for crRNA pairing. [Figure 28]This diagram depicts the single guide RNA folding of the active MG64-1, MG64-2, and MG64-6 CAST systems. An active, manipulated sgRNA of MG64-1 is also shown. [Figure 29A] This diagram illustrates in vitro screening for CAST transpositions using a PAM library. Figure 29A depicts the screening setup for in vitro PAM determination. [Figure 29B] In vitro screening for CAST rearrangements using a PAM library is illustrated. Figure 29B shows a schematic diagram of junctional PCR for detecting rearrangement products. [Figure 30A] This diagram depicts the dislocation junctions of MG64-1 CAST (left lane) and MG64-6 CAST (right lane) amplified by PCR. [Figure 30B] The SeqLogo representation of the PAM detected by MG64-1 is shown (above). [Figure 30C] The combined frequency of MG64-1 is plotted against its proximal and distal distances. [Figure 31] The single-guide RNA engineering of 64-1 is depicted. Deletion of a region of approximately 130 bp–190 bp (green and blue-green sections of the structure) generated an sgRNA that induced a potent rearrangement reaction (green bars on the heatmap). [Figure 32] This study describes the cross-reactivity of MG64-2 sgRNA with PAM in combinations of MG64-1 and MG64-2 sgRNA + MG64-1 effector. [Figure 33-1] The sequence diagram and secondary structure prediction model depict single guide RNA cleavage in the coding DNA of MG64-2 sgRNA. Deletion regions and cleavage in the sequence diagram are shown as bars (del1, del2, del3, del4, del5, and del6). In the secondary structure prediction model, deletion regions within ellipses indicate tested cleavage (pseudoknots, deletions 1, 2, 3, and 4 across 5 and 6). [Figure 33-2]The sequence diagram and secondary structure prediction model depict single guide RNA cleavage in the coding DNA of MG64-2 sgRNA. Deletion regions and cleavage in the sequence diagram are shown as bars (del1, del2, del3, del4, del5, and del6). In the secondary structure prediction model, deletion regions within ellipses indicate tested cleavage (pseudoknots, deletions 1, 2, 3, and 4 across 5 and 6). [Figure 34] The data illustrates the activity of the manipulated MG64-2 sgRNA in the MG64-1 CAST system. PCR reactions represent each possible integration conjugate or negative control (Panel B in Figure 29). Products of successful integration are highlighted with arrows. Lanes enclosed in boxes are not relevant to this experiment. [Figure 35] This describes the design of the MG64-2 sgRNA splitting guide. sgRNA fragments were synthesized separately and then re-annealed before being tested in transposition experiments. [Figure 36] This diagram illustrates data demonstrating the activity of split MG64-2 sgRNA in the MG64-1 CAST system. PCR reactions represent each possible fusion conjugation or negative control (Panel B, Figure 29). Products of successful fusion are highlighted with arrows. Lanes enclosed in boxes are irrelevant to this experiment. [Figure 37] The data illustrates that minimizing LE and RE maintained the rearrangement activity of this system. [Figure 38A] The results of E. coli integration with MG64-1 are depicted. Figure 38A shows a schematic diagram of the CAST system integration into E. coli. Figure 38B shows NGS data demonstrating editing efficiency of over 80%. Figure 38C shows off-target analysis showing that no off-target integration was detected for more than 1% of all total dislocation events. [Figure 38B] The results of E. coli integration with MG64-1 are depicted. Figure 38A shows a schematic diagram of the CAST system integration into E. coli. Figure 38B shows NGS data demonstrating editing efficiency of over 80%. Figure 38C shows off-target analysis showing that no off-target integration was detected for more than 1% of all total dislocation events. [Figure 38C] The results of E. coli integration with MG64-1 are depicted. Figure 38A shows a schematic diagram of the CAST system integration into E. coli. Figure 38B shows NGS data demonstrating editing efficiency of over 80%. Figure 38C shows off-target analysis showing that no off-target integration was detected for more than 1% of all total dislocation events. [Figure 39] This study depicts the local insertion rates of various endogenous loci in the E. coli genome. [Figure 40A] The results of multi-gene coordinate systemization are depicted. Figure 40A depicts the local insertion frequencies at endogenous and engineered loci, respectively. Figure 40B depicts the relative insertion frequencies of on-target insertions at endogenous loci, on-target insertions at engineered loci, and off-target insertions. Integration at both combined loci accounted for over 95% of all integrations occurring on the genome. [Figure 40B] The results of multi-gene coordinate systemization are depicted. Figure 40A depicts the local insertion frequencies at endogenous and engineered loci, respectively. Figure 40B depicts the relative insertion frequencies of on-target insertions at endogenous loci, on-target insertions at engineered loci, and off-target insertions. Integration at both combined loci accounted for over 95% of all integrations occurring on the genome. [Figure 41] This diagram illustrates Sanger sequencing data of an integrated PCR product demonstrating the in vitro activity of MG64-1. The reaction is that of the RE donor target product, and the point at which sequencing stops matching the donor DNA is when a junction is formed (dark bar below the sequencing peak). [Figure 42A] A schematic diagram of serial dilution of target DNA for in vitro transposition experiments is shown. CAST components are expressed and added to the reaction with sgRNA and donor plasmid transcribed in vitro. Target plasmid DNA is added at decreasing concentrations and tested for transposition experiments. Once the minimum amount of target DNA is determined, the transposition reaction is assayed by adding increasing amounts of human genomic DNA. [Figure 42B]The diagram shows the PCR amplification of the rearrangement reaction. The 8N PAM plasmid library (8N-target, Rxn#1) is targeted in the CAST system to integrate with donor DNA (Rxn#2). If integration is successful, a junction PCR reaction is performed using primers to amplify four putative integration reactions (Rxn#3, #4, #5, and #6) based on the orientation of cargo integration. [Figure 42C] The PCR reaction products from an in vitro transposition assay using serial dilutions of target plasmid DNA are illustrated. The target, donor, and reactions #3, #4, #5, and #6 correspond to the combined PCR products shown in Figure 42B. [Figure 42D] The PCR reaction products from an in vitro transposition assay using a fixed amount of target plasmid DNA (0.5 ng) are shown, with the search space increased by adding an increasing amount of human genomic DNA. The target, donor, and reactions #3, #4, #5, and #6 correspond to the combined PCR products shown in Figure 42B. [Figure 43A] A schematic diagram of the transposition reaction across high-copy elements is shown. The target PCR product extends to the wild-type target element when assayed using the CAST protein and an sgRNA targeting one of several sequenced targets. Incorporation can occur in either forward orientation, reverse orientation, or both. Forward transposition products are assayed by junction PCR (Fwd PCR), which amplifies the region containing the LE of the donor DNA up to the 5' end of the target site. Reverse junction reactions assay the region containing the LE of the donor DNA up to the 3' end of the target element (Rev PCR). [Figure 43B] This figure shows PCR reaction products from in vitro transposition assays at 15 target sites (guides) in the line 1 3' element of human genomic DNA. The forward PCR and rev PCR of the targets and reactions correspond to the combined PCR products shown in Figure 43A. [Figure 43C]This figure shows PCR reaction products from in vitro transposition assays at 15 target sites (guides) in SVA elements of human genomic DNA. The forward and reverse PCR results for the targets and reactions correspond to the PCR integration products shown in Figure 43A. Bands highlighted with arrows indicate successful target integration. [Figure 43D] This figure shows PCR reaction products from in vitro transposition assays at 15 target sites (guides) in the HERV element of human genomic DNA. The forward and reverse PCR of the targets and reactions correspond to the PCR integration products shown in Figure 43A. Bands highlighted with arrows indicate successful target integration. [Figure 43E] Line 1 shows Sanger sequencing of Fwd PCR integration products at multiple target sites of the 3' element. Integration occurs where the sequencing trace stops matching the donor DNA (gray vertical bars). [Figure 43F] Line 1 shows Sanger sequencing of Rev PCR integration products at multiple target sites in the 3' element. Integration occurs where the sequencing trace stops matching the target DNA (gray vertical bars). [Figure 43G] This shows the Sanger sequencing of the Fwd PCR integration product at SVA target site 3. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. [Figure 43H] This shows the Sanger sequencing of the Fwd PCR product at HERV target site 5. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. [Figure 44] This figure shows PCR reaction products from in vitro transposition assays at line 1 target sites 12 and 15 in human genomic DNA containing functional domains. The forward and reverse PCR of the target and reaction correspond to the PCR integration products shown in Figure 42A. Bands highlighted with arrows indicate successful target integration. [Figure 45A]In vitro rearrangement experiments using CAST, S15, NLS-S15, and S15-NLS expressed from eukaryotic transcription / translation reactions are illustrated. Figure 45A shows the in vitro rearrangement reactions of MG64-1 with CAST and S15. The wheat germ extract-expressed CAST component promotes rearrangement, albeit at a low rate, without the addition of S15 (faint band highlighted by arrows). The addition of PURExpress reagent (used PUREx) increases the rearrangement efficiency, as indicated by the band intensity in Rxn#5 (PURExpress reagent contains S15). The independent addition of S15 and S15-NLS translated from the wheat germ extract reaction increases the rearrangement efficiency by MG64-1 in vitro compared to other conditions tested (strong band highlighted by arrows). Figure 45B shows the in vitro reaction of rearrangement using the NLS-S15 configuration. The addition of PURExpress reagent increased in vitro rearrangement compared to the condition with only the CAST component (lane 2) (lane 3). The NLS-S15 configuration did not improve rearrangement (lanes 4-5). Rxn#5 enclosed in a box represents the expected band when rearrangement activity is detected. [Figure 45B]In vitro rearrangement experiments using CAST, S15, NLS-S15, and S15-NLS expressed from eukaryotic transcription / translation reactions are illustrated. Figure 45A shows the in vitro rearrangement reactions of MG64-1 with CAST and S15. The wheat germ extract-expressed CAST component promotes rearrangement, albeit at a low rate, without the addition of S15 (faint band highlighted by arrows). The addition of PURExpress reagent (used PUREx) increases the rearrangement efficiency, as indicated by the band intensity in Rxn#5 (PURExpress reagent contains S15). The independent addition of S15 and S15-NLS translated from the wheat germ extract reaction increases the rearrangement efficiency by MG64-1 in vitro compared to other conditions tested (strong band highlighted by arrows). Figure 45B shows the in vitro reaction of rearrangement using the NLS-S15 configuration. The addition of PURExpress reagent increased in vitro rearrangement compared to the condition with only the CAST component (lane 2) (lane 3). The NLS-S15 configuration did not improve rearrangement (lanes 4-5). Rxn#5 enclosed in a box represents the expected band when rearrangement activity is detected. [Figure 46A]A schematic diagram of the fusion plasmid for cell translocation is shown. Figure 46A: Two targeting complex plasmids and one donor plasmid are assembled for high-copy element lines 1, targets 8, 12, and 15, and SVA target 3. Figure 46B shows cell translocation to high-copy elements using H1 core-TniQ or HMGN1-TniQ at line 1 targets 8, 12, 15, and SVA target 3. Arrows indicate amplified translocation junction reactions in either forward orientation (Fwd PCR) or reverse orientation (Rev PCR). The simulated control represents the reaction without targeting plasmids or donor plasmids. Figure 46C shows Sanger sequencing of the Fwd PCR of the PCR fusion product at line 1 3' target site 8. Fusion was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where fusion occurs. Figure 46D shows the Sanger sequencing of the PCR-integrated product Fwd PCR at target site 8 of line 1 3'. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46E shows the Sanger sequencing of the PCR-integrated product Rev PCR at target site 12 of line 1 3'. Integration was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46F shows the Sanger sequencing of the PCR-integrated product Rev PCR at target site 12 of line 1 3'. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46G shows the Sanger sequencing of the Fwd PCR of the PCR integration product at the 3' target site 15 of line 1. Integration was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where integration occurs.Figure 46H shows the Sanger sequencing of the Fwd PCR of the PCR integration product at target site 15 of line 1. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where integration occurs. [Figure 46B]A schematic diagram of the fusion plasmid for cell translocation is shown. Figure 46A: Two targeting complex plasmids and one donor plasmid are assembled for high-copy element lines 1, targets 8, 12, and 15, and SVA target 3. Figure 46B shows cell translocation to high-copy elements using H1 core-TniQ or HMGN1-TniQ at line 1 targets 8, 12, 15, and SVA target 3. Arrows indicate amplified translocation junction reactions in either forward orientation (Fwd PCR) or reverse orientation (Rev PCR). The simulated control represents the reaction without targeting plasmids or donor plasmids. Figure 46C shows Sanger sequencing of the Fwd PCR of the PCR fusion product at line 1 3' target site 8. Fusion was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where fusion occurs. Figure 46D shows the Sanger sequencing of the PCR-integrated product Fwd PCR at target site 8 of line 1 3'. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46E shows the Sanger sequencing of the PCR-integrated product Rev PCR at target site 12 of line 1 3'. Integration was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46F shows the Sanger sequencing of the PCR-integrated product Rev PCR at target site 12 of line 1 3'. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46G shows the Sanger sequencing of the Fwd PCR of the PCR integration product at the 3' target site 15 of line 1. Integration was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where integration occurs.Figure 46H shows the Sanger sequencing of the Fwd PCR of the PCR integration product at target site 15 of line 1. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where integration occurs. [Figure 46C]A schematic diagram of the fusion plasmid for cell translocation is shown. Figure 46A: Two targeting complex plasmids and one donor plasmid are assembled for high-copy element lines 1, targets 8, 12, and 15, and SVA target 3. Figure 46B shows cell translocation to high-copy elements using H1 core-TniQ or HMGN1-TniQ at line 1 targets 8, 12, 15, and SVA target 3. Arrows indicate amplified translocation junction reactions in either forward orientation (Fwd PCR) or reverse orientation (Rev PCR). The simulated control represents the reaction without targeting plasmids or donor plasmids. Figure 46C shows Sanger sequencing of the Fwd PCR of the PCR fusion product at line 1 3' target site 8. Fusion was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where fusion occurs. Figure 46D shows the Sanger sequencing of the PCR-integrated product Fwd PCR at target site 8 of line 1 3'. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46E shows the Sanger sequencing of the PCR-integrated product Rev PCR at target site 12 of line 1 3'. Integration was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46F shows the Sanger sequencing of the PCR-integrated product Rev PCR at target site 12 of line 1 3'. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46G shows the Sanger sequencing of the Fwd PCR of the PCR integration product at the 3' target site 15 of line 1. Integration was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where integration occurs.Figure 46H shows the Sanger sequencing of the Fwd PCR of the PCR integration product at target site 15 of line 1. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where integration occurs. [Figure 46D]A schematic diagram of the fusion plasmid for cell translocation is shown. Figure 46A: Two targeting complex plasmids and one donor plasmid are assembled for high-copy element lines 1, targets 8, 12, and 15, and SVA target 3. Figure 46B shows cell translocation to high-copy elements using H1 core-TniQ or HMGN1-TniQ at line 1 targets 8, 12, 15, and SVA target 3. Arrows indicate amplified translocation junction reactions in either forward orientation (Fwd PCR) or reverse orientation (Rev PCR). The simulated control represents the reaction without targeting plasmids or donor plasmids. Figure 46C shows Sanger sequencing of the Fwd PCR of the PCR fusion product at line 1 3' target site 8. Fusion was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where fusion occurs. Figure 46D shows the Sanger sequencing of the PCR-integrated product Fwd PCR at target site 8 of line 1 3'. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46E shows the Sanger sequencing of the PCR-integrated product Rev PCR at target site 12 of line 1 3'. Integration was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46F shows the Sanger sequencing of the PCR-integrated product Rev PCR at target site 12 of line 1 3'. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46G shows the Sanger sequencing of the Fwd PCR of the PCR integration product at the 3' target site 15 of line 1. Integration was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where integration occurs.Figure 46H shows the Sanger sequencing of the Fwd PCR of the PCR integration product at target site 15 of line 1. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where integration occurs. [Figure 46E]A schematic diagram of the fusion plasmid for cell translocation is shown. Figure 46A: Two targeting complex plasmids and one donor plasmid are assembled for high-copy element lines 1, targets 8, 12, and 15, and SVA target 3. Figure 46B shows cell translocation to high-copy elements using H1 core-TniQ or HMGN1-TniQ at line 1 targets 8, 12, 15, and SVA target 3. Arrows indicate amplified translocation junction reactions in either forward orientation (Fwd PCR) or reverse orientation (Rev PCR). The simulated control represents the reaction without targeting plasmids or donor plasmids. Figure 46C shows Sanger sequencing of the Fwd PCR of the PCR fusion product at line 1 3' target site 8. Fusion was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where fusion occurs. Figure 46D shows the Sanger sequencing of the PCR-integrated product Fwd PCR at target site 8 of line 1 3'. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46E shows the Sanger sequencing of the PCR-integrated product Rev PCR at target site 12 of line 1 3'. Integration was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46F shows the Sanger sequencing of the PCR-integrated product Rev PCR at target site 12 of line 1 3'. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46G shows the Sanger sequencing of the Fwd PCR of the PCR integration product at the 3' target site 15 of line 1. Integration was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where integration occurs.Figure 46H shows the Sanger sequencing of the Fwd PCR of the PCR integration product at target site 15 of line 1. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where integration occurs. [Figure 46F]A schematic diagram of the fusion plasmid for cell translocation is shown. Figure 46A: Two targeting complex plasmids and one donor plasmid are assembled for high-copy element lines 1, targets 8, 12, and 15, and SVA target 3. Figure 46B shows cell translocation to high-copy elements using H1 core-TniQ or HMGN1-TniQ at line 1 targets 8, 12, 15, and SVA target 3. Arrows indicate amplified translocation junction reactions in either forward orientation (Fwd PCR) or reverse orientation (Rev PCR). The simulated control represents the reaction without targeting plasmids or donor plasmids. Figure 46C shows Sanger sequencing of the Fwd PCR of the PCR fusion product at line 1 3' target site 8. Fusion was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where fusion occurs. Figure 46D shows the Sanger sequencing of the PCR-integrated product Fwd PCR at target site 8 of line 1 3'. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46E shows the Sanger sequencing of the PCR-integrated product Rev PCR at target site 12 of line 1 3'. Integration was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46F shows the Sanger sequencing of the PCR-integrated product Rev PCR at target site 12 of line 1 3'. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46G shows the Sanger sequencing of the Fwd PCR of the PCR integration product at the 3' target site 15 of line 1. Integration was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where integration occurs.Figure 46H shows the Sanger sequencing of the Fwd PCR of the PCR integration product at target site 15 of line 1. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where integration occurs. [Figure 46G]A schematic diagram of the fusion plasmid for cell translocation is shown. Figure 46A: Two targeting complex plasmids and one donor plasmid are assembled for high-copy element lines 1, targets 8, 12, and 15, and SVA target 3. Figure 46B shows cell translocation to high-copy elements using H1 core-TniQ or HMGN1-TniQ at line 1 targets 8, 12, 15, and SVA target 3. Arrows indicate amplified translocation junction reactions in either forward orientation (Fwd PCR) or reverse orientation (Rev PCR). The simulated control represents the reaction without targeting plasmids or donor plasmids. Figure 46C shows Sanger sequencing of the Fwd PCR of the PCR fusion product at line 1 3' target site 8. Fusion was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where fusion occurs. Figure 46D shows the Sanger sequencing of the PCR-integrated product Fwd PCR at target site 8 of line 1 3'. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46E shows the Sanger sequencing of the PCR-integrated product Rev PCR at target site 12 of line 1 3'. Integration was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46F shows the Sanger sequencing of the PCR-integrated product Rev PCR at target site 12 of line 1 3'. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46G shows the Sanger sequencing of the Fwd PCR of the PCR integration product at the 3' target site 15 of line 1. Integration was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where integration occurs.Figure 46H shows the Sanger sequencing of the Fwd PCR of the PCR integration product at target site 15 of line 1. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where integration occurs. [Figure 46H]A schematic diagram of the fusion plasmid for cell translocation is shown. Figure 46A: Two targeting complex plasmids and one donor plasmid are assembled for high-copy element lines 1, targets 8, 12, and 15, and SVA target 3. Figure 46B shows cell translocation to high-copy elements using H1 core-TniQ or HMGN1-TniQ at line 1 targets 8, 12, 15, and SVA target 3. Arrows indicate amplified translocation junction reactions in either forward orientation (Fwd PCR) or reverse orientation (Rev PCR). The simulated control represents the reaction without targeting plasmids or donor plasmids. Figure 46C shows Sanger sequencing of the Fwd PCR of the PCR fusion product at line 1 3' target site 8. Fusion was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where fusion occurs. Figure 46D shows the Sanger sequencing of the PCR-integrated product Fwd PCR at target site 8 of line 1 3'. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46E shows the Sanger sequencing of the PCR-integrated product Rev PCR at target site 12 of line 1 3'. Integration was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46F shows the Sanger sequencing of the PCR-integrated product Rev PCR at target site 12 of line 1 3'. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) indicates where integration occurs. Figure 46G shows the Sanger sequencing of the Fwd PCR of the PCR integration product at the 3' target site 15 of line 1. Integration was mediated by MG64-1 with an NLS-H1 core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where integration occurs.Figure 46H shows the Sanger sequencing of the Fwd PCR of the PCR integration product at target site 15 of line 1. Integration was mediated by MG64-1 with an NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is where integration occurs. [Figure 47A] This shows immunofluorescence staining for the localization of Cas12k CAST components in human cells. Figure 47A: Top row: Detection of TnsB localization, Middle row: Detection of Cas12k localization, Bottom row: Detection of TnsC localization. The images show that MG64-1 Cas12k and TnsB localize in the nucleus of mammalian cells, while TnsC localizes in the cytoplasm. The Cas12k CAST protein was tagged with an HA tag. An anti-HA antibody was used for protein detection. DAPI was used to stain the DNA (nucleus). Figure 47B: Top and bottom rows: Detection of TniQ localization. The images show that MG64-1 TniQ localizes in the nucleus of mammalian cells. The CAST protein was tagged with an HA tag. An anti-HA antibody was used for protein detection. DAPI was used to stain the DNA (nucleus). Figure 47C: All rows: Detection of TnsC co-localization with TniQ. The image shows that while some TnsCs may remain in the cytoplasm, here they co-localize with TniQ in the nucleus. The CAST protein was tagged with an HA tag. An anti-HA antibody was used for protein detection. DAPI was used to stain the DNA (nucleus). Figure 47D: Both columns: Cas12k, TnsB, TnsC, and TniQ, co-delivered to HEK293T cells, localize in the nucleus. The CAST protein was tagged with an HA tag. An anti-HA antibody was used for protein detection. DAPI was used to stain the DNA (nucleus). [Figure 47B]This shows immunofluorescence staining for the localization of Cas12k CAST components in human cells. Figure 47A: Top row: Detection of TnsB localization, Middle row: Detection of Cas12k localization, Bottom row: Detection of TnsC localization. The images show that MG64-1 Cas12k and TnsB localize in the nucleus of mammalian cells, while TnsC localizes in the cytoplasm. The Cas12k CAST protein was tagged with an HA tag. An anti-HA antibody was used for protein detection. DAPI was used to stain the DNA (nucleus). Figure 47B: Top and bottom rows: Detection of TniQ localization. The images show that MG64-1 TniQ localizes in the nucleus of mammalian cells. The CAST protein was tagged with an HA tag. An anti-HA antibody was used for protein detection. DAPI was used to stain the DNA (nucleus). Figure 47C: All rows: Detection of TnsC co-localization with TniQ. The image shows that while some TnsCs may remain in the cytoplasm, here they co-localize with TniQ in the nucleus. The CAST protein was tagged with an HA tag. An anti-HA antibody was used for protein detection. DAPI was used to stain the DNA (nucleus). Figure 47D: Both columns: Cas12k, TnsB, TnsC, and TniQ, co-delivered to HEK293T cells, localize in the nucleus. The CAST protein was tagged with an HA tag. An anti-HA antibody was used for protein detection. DAPI was used to stain the DNA (nucleus). [Figure 47C]This shows immunofluorescence staining for the localization of Cas12k CAST components in human cells. Figure 47A: Top row: Detection of TnsB localization, Middle row: Detection of Cas12k localization, Bottom row: Detection of TnsC localization. The images show that MG64-1 Cas12k and TnsB localize in the nucleus of mammalian cells, while TnsC localizes in the cytoplasm. The Cas12k CAST protein was tagged with an HA tag. An anti-HA antibody was used for protein detection. DAPI was used to stain the DNA (nucleus). Figure 47B: Top and bottom rows: Detection of TniQ localization. The images show that MG64-1 TniQ localizes in the nucleus of mammalian cells. The CAST protein was tagged with an HA tag. An anti-HA antibody was used for protein detection. DAPI was used to stain the DNA (nucleus). Figure 47C: All rows: Detection of TnsC co-localization with TniQ. The image shows that while some TnsCs may remain in the cytoplasm, here they co-localize with TniQ in the nucleus. The CAST protein was tagged with an HA tag. An anti-HA antibody was used for protein detection. DAPI was used to stain the DNA (nucleus). Figure 47D: Both columns: Cas12k, TnsB, TnsC, and TniQ, co-delivered to HEK293T cells, localize in the nucleus. The CAST protein was tagged with an HA tag. An anti-HA antibody was used for protein detection. DAPI was used to stain the DNA (nucleus). [Figure 47D]This shows immunofluorescence staining for the localization of Cas12k CAST components in human cells. Figure 47A: Top row: Detection of TnsB localization, Middle row: Detection of Cas12k localization, Bottom row: Detection of TnsC localization. The images show that MG64-1 Cas12k and TnsB localize in the nucleus of mammalian cells, while TnsC localizes in the cytoplasm. The Cas12k CAST protein was tagged with an HA tag. An anti-HA antibody was used for protein detection. DAPI was used to stain the DNA (nucleus). Figure 47B: Top and bottom rows: Detection of TniQ localization. The images show that MG64-1 TniQ localizes in the nucleus of mammalian cells. The CAST protein was tagged with an HA tag. An anti-HA antibody was used for protein detection. DAPI was used to stain the DNA (nucleus). Figure 47C: All rows: Detection of TnsC co-localization with TniQ. The image shows that while some TnsCs may remain in the cytoplasm, here they co-localize with TniQ in the nucleus. The CAST protein was tagged with an HA tag. An anti-HA antibody was used for protein detection. DAPI was used to stain the DNA (nucleus). Figure 47D: Both columns: Cas12k, TnsB, TnsC, and TniQ, co-delivered to HEK293T cells, localize in the nucleus. The CAST protein was tagged with an HA tag. An anti-HA antibody was used for protein detection. DAPI was used to stain the DNA (nucleus). [Figure 48A] This describes the in vitro screening of the MG64-1 Cas12k CAST rearrangement. Figure 48A: Schematic diagram of the construct used for MG64-1 holocomplex purification. Figure 48B: Schematic diagram of junctional PCR for detection of the rearrangement product. A 5'PAM, followed by a target substrate with a protospacer (target, Rxn#1), is targeted in the CAST system to incorporate cargo DNA (Rxn#2). If the incorporation is successful, a junctional PCR reaction is performed using primers to amplify four putative incorporation reactions based on the orientation of the cargo incorporation. [Figure 48B]This describes the in vitro screening of the MG64-1 Cas12k CAST rearrangement. Figure 48A: Schematic diagram of the construct used for MG64-1 holocomplex purification. Figure 48B: Schematic diagram of junctional PCR for detection of the rearrangement product. A 5'PAM, followed by a target substrate with a protospacer (target, Rxn#1), is targeted in the CAST system to incorporate cargo DNA (Rxn#2). If the incorporation is successful, a junctional PCR reaction is performed using primers to amplify four putative incorporation reactions based on the orientation of the cargo incorporation. [Figure 48C] Figure 48C: Fractions collected during 2L scale purification of MG64-1 holocomplex electrophoresis on an unstained, denatured PAGE gel. Figure 48D: Chromatogram of size exclusion chromatography (SEC) performed on the MG64-1 holocomplex. The peak centered at 29.3 mL (peak 1) was used for the in vitro activity assay. [Figure 48D] Figure 48C: Fractions collected during 2L scale purification of MG64-1 holocomplex electrophoresis on an unstained, denatured PAGE gel. Figure 48D: Chromatogram of size exclusion chromatography (SEC) performed on the MG64-1 holocomplex. The peak centered at 29.3 mL (peak 1) was used for the in vitro activity assay. [Figure 49A] This describes in vitro transposition using the Peak 1 recovered holocomplex supplemented with TnT expression components. Lane L) Ladder, Lane 1) TnT expression CAST component apo condition (-sgRNA), Lane 2) TnT expression CAST component holo condition (+sgRNA), Lane 3) Purified Peak 1 supplemented with TnT CAST component without additional supplementation of Cas12k (-TnT Cas12k), Lane 4) Purified Peak 1 supplemented with TnT CAST component without additional supplementation of TnsC (-TnT TnsC), Lane 5) Peak 1 supplemented with TnT CAST component without additional supplementation of TniQ (-TnT TniQ), Lane 6) Peak 1 supplemented with TnT CAST component without additional supplementation of S15 (-TnT S15). [Figure 49B]Figure 49B depicts the Sanger sequencing of lanes 3, 4, 5, and 6 from both the pDonor and target directions of the amplified LE-PAM target-donor junction. The vertical line defines the predicted translocation junction for MG64-1 in the reference sequence. Degradation of the signal from either direction results from a multitude of signals reflected in the PCR amplification. [Figure 50] This diagram illustrates the identification of ribosomal protein S15 homologs in cyanobacterial genome fragments. Candidate sequences from the same sample from which MG64-1 was recovered are highlighted with dark circles. Reference S15 from E. coli is indicated by an arrow. [Figure 51A] Figures 51A and 51B show schematic diagrams of the dual transcription component (Figure 51A) and the all-in-one transposition component (Figure 51B). In the dual transcription system, one transcript is under the control of the CMV-betaglobin promoter on a single plasmid. Cas12k-sso7d-NLS is ligated to S15 via a 2A self-cleaving peptide, and the IRES element isolates the second ORF, the NLS functional domain-TniQ, where the functional domain is H1-core or HMGN1. The plasmid also contains either a non-targeted (null) MG64-1 single guide or a targeted MG64-1 single guide. In the second plasmid, the CMV-betaglobin promoter drives the transcription of NLS-TnsB and NLS-TnsC isolated by the IRES element, and TIRs (LE and RE) are adjacent to bacterial replication origin and antibiotic resistance markers. Under single transcription conditions (Figure 51B), a single CMV-betaglobin promoter controls the expression of an all-in-one transcript, which consists of Cas12k-sso7d-NLS, S15-NLS-functional domain-TniQ, NLS-TnsB, and NLS-TnsC, separated at the 2A and IRES elements. This single helper plasmid (pHelper) also contains either a single guide to non-targeted or targeted MG64-1 under the control of the pU6 promoter. A second plasmid is a p-donor with LE and RE flanking either a reporter gene, therapeutic transcript, or a select marker. [Figure 51B] Figures 51A and 51B show schematic diagrams of the dual transcription component (Figure 51A) and the all-in-one transposition component (Figure 51B). In the dual transcription system, one transcript is under the control of the CMV-betaglobin promoter on a single plasmid. Cas12k-sso7d-NLS is ligated to S15 via a 2A self-cleaving peptide, and the IRES element isolates the second ORF, the NLS functional domain-TniQ, where the functional domain is H1-core or HMGN1. The plasmid also contains either a non-targeted (null) MG64-1 single guide or a targeted MG64-1 single guide. In the second plasmid, the CMV-betaglobin promoter drives the transcription of NLS-TnsB and NLS-TnsC isolated by the IRES element, and TIRs (LE and RE) are adjacent to bacterial replication origin and antibiotic resistance markers. Under single transcription conditions (Figure 51B), a single CMV-betaglobin promoter controls the expression of an all-in-one transcript, which consists of Cas12k-sso7d-NLS, S15-NLS-functional domain-TniQ, NLS-TnsB, and NLS-TnsC, separated at the 2A and IRES elements. This single helper plasmid (pHelper) also contains either a single guide to non-targeted or targeted MG64-1 under the control of the pU6 promoter. A second plasmid is a p-donor with LE and RE flanking either a reporter gene, therapeutic transcript, or a select marker. [Figure 52]Figure 52 shows the testing of the all-in-one pHelper H1 core plasmid in human cells. All transpositions are shown as junction bands in the forward LE image. Lane 1: Testing of a dual transcription system with a non-targeted single guide (Figure 51A). Lane 2: Testing of a dual transcription system encoding a targeted single guide for integration to the 3' target 8 of line 1. Lanes 3-5: Testing of single transcript pHelper systems with pDonors expressing the fluorescent protein mNeon. Lane 3: All-in-one pHelper and mNeon pDonor without a targeted single guide. Lane 4: All-in-one pHelper and mNeon pDonor with a guide targeting the 3' target 8 of line 1 in a ratio of 6 μg pHelper to 12 μg pDonor. Lane 5 is an all-in-one pHelper and mNeon pDonor with a guide targeting the 3' target 8 of line 1 in a ratio of 12 μg pHelper to 6 μg pDonor. [Figure 53] Figure 53 shows the testing of non-replicating donors compared to replicating donors. Non-replicating donors are cleaved at the SV40 origin element, which allows the plasmid to proliferate in human cells. Integration is performed using an all-in-one vector. Lane 1: Simulated transfection control. Lane 2: All-in-one pHelper targeting 3' target 8 of Line 1 with a replicating donor as a positive control. Lanes 3-5: All-in-one pHelper without a targeted single guide, with a replicating mNeon donor. Lanes 6-8: All-in-one pHelper with a single guide at 3' target 8 of Line 1 with a replicating mNeon donor. Lanes 9-11: All-in-one pHelper without a targeted single guide, with a non-replicating mNeon donor. Lanes 12-14: All-in-one pHelper with a single guide at 3' target 8 of Line 1 with a non-replicating mNeon donor. [Figure 54]Figure 54 shows the addition of Clpx to transduction in cells. Lane 1: All-in-one pHelper with a single non-targeting guide using a replicated mNeon pDonor. Lane 2: All-in-one targeting targeting 3' target 8 of Line 1 with a replicated mNeon donor as a positive control. Lanes 3-6: ClpX-NLS plasmid added to the transfer plasmid mixture in amounts of (0.25 μg, 0.5 μg, 1 μg, and 2 μg) under the same conditions as Lane 2. Lanes 7-10: NLS-ClpX plasmid added to the transfer plasmid mixture in amounts of (0.25 μg, 0.5 μg, 1 μg, and 2 μg) under the same conditions as Lane 2. [Figure 55] Figure 55 shows a schematic diagram of a three-plasmid transfection system targeting a single-copy locus. The pHelper plasmid contains a single guide targeting the desired single-copy locus. A replica plasmid is used for all plasmids: pHelper, pDonor, and pClpX. pDonor has a structure in which a TIR sequence flanks a pCMV-BG promoter that drives the expression of the mNeon fluorescent protein. pClpX is a pCMV-betaglobin promoter that drives the E. coli ClpX-NLS sequence. All three plasmids are transfected into HEK293T cells in the following ratios: 12 μg pHelper: 6 μg pDonor: 1 μg pClpX: 54 μL Mirus-LT1 (transfection reagent). The cells are then incubated at 37°C for 72 hours and harvested for their gDNA. [Figure 56] Figure 56 shows the transposition of a pDonor to the single-copy target AAVS1. AAVS1 is a safe harbor locus that enables non-deletion expression of exogenous genes in human cells when the cassette is integrated into the human genome. AAVS1 is expressed only once in the human genome on chromosome 19. By assaying junction PCR of genomic DNA taken from cells transfected with pHelper targeting both AAVS1 target 5 and AAVS1 target 6, it was possible to visualize the transposition of the pDonor in the forward LE direction to its intended target. [Figure 57]Figure 57 shows the Sanger sequencing of the dislocated AAVS1 targets 5 and 6 in the anterior LE direction. The dislocation of AAVS5 has a primary dislocation 63 bp away from the PAM, and the dislocation of AAVS1 target 6 has a primary transposition event 61 bp away from the PAM. [Figure 58A] This shows the metastatic NGS sequencing at AAVS target 5. [Figure 58B] This shows the metastatic NGS sequencing at AAVS target 6. [Figure 59] Figure 59 shows a schematic diagram of NGS quantification. Hypothetical transpositions are modeled on unintegrated target sequences. The two primers should be as close together as possible. [Figure 60] Figure 60 shows transposition at AAVS1 target 5. A 6g DNA preparation of AAVS1 target 5 pHelper was amplified with transposition-specific primers (upper panel) that have a transposition band reflecting the posterior LE junction reaction from the genomic target to the donor cargo. Target pHelper is not a culture where the target is not identified relative to the spacer sequence. Without the target pHelper condition, the transposition band is not visible. Lower panel: 25x cycle PCR amplification shows equal loadings for the reaction to the indexing step for NGS sequencing for both AAVS1 target 5 pHelper-treated samples and samples without target pHelper treatment. [Figure 61A] Figures 61A to 61C show NGS efficiency and allelic visualization of unedited alleles. Figure 61A is a graph showing the measurement of NGS efficiency of AAVS1 target 5 compared to the control. [Figure 61B] Figures 61A–61C show NGS efficiency and allele visualization of unedited alleles. Figure 61B shows allele mapping of unedited alleles in targeted NGS readouts, demonstrating that only a small percentage of readouts can be mapped to a reference using SNPs. [Figure 61C]Figures 61A–61C show NGS efficiency and allele visualization of unedited alleles. Figure 61C shows allele mapping of unedited alleles in AAVS1 target 5 readouts, showing that SNPs have a similar rate of readouts, indicating that they occurred at genomic loci not attributable to AAVS1 target 5 integration on pHelper. [Figure 62] Figure 62 shows a schematic diagram of the plasmid used in the experiments of this disclosure. [Figure 63A] Figures 63A and 63B show the integration of linearized donors in HEK293T cells. Figure 63A shows NGS integration of a single guideless (pHelper Null) or AAVS1-5 targeted (pHelper AIO) linear CAST pDonor. Figure 63B shows ddPCR integration of a single guideless (pHelper Null) or AAVS1-5 targeted (pHelper AIO) linear CAST pDonor. [Figure 63B] Figures 63A and 63B show the integration of linearized donors in HEK293T cells. Figure 63A shows NGS integration of a single guideless (pHelper Null) or AAVS1-5 targeted (pHelper AIO) linear CAST pDonor. Figure 63B shows ddPCR integration of a single guideless (pHelper Null) or AAVS1-5 targeted (pHelper AIO) linear CAST pDonor. [Figure 64] Figure 64 shows plasmid donor integration in K562 cells. Integration of CAST pDonor and single guide expression cassette by NGS: Nucleofection using a single guide targeting AAVS1-5 and without helper plasmids (pHelper Null) or a single guide containing AAVS1-5 targets (pHelper AIO). Includes 2 × 10⁵ K562 cells (left) or 5 × 10⁵ K562 cells (right). [Figure 65A] Figures 65A-65B show the translocation of Hep3B using lipofectamine 2000 and lipofectamine 3000, and the NGS quantification of p donors containing sg cassettes. [Figure 65B] Figures 65A-65B show the translocation of Hep3B using lipofectamine 2000 and lipofectamine 3000, and the NGS quantification of p donors containing sg cassettes. [Figure 66] Figure 66 shows NGS detection of transpositions in human albumin (ALB) in HEK293T. Guide targets 1077–1138 were predicted sites for albumin intron 1. 1139 and 1140 were guides constructed for AAVS1. [Figure 67A] Figures 67A to 67C show the NGS read alignments of modeled translocation targets located 60 bp away from the PAM for targets 1093 (Figure 67A (sequence numbers 1264 to 1274 in order of appearance)), 1101 (Figure 67B (sequence numbers 1275 to 1279 in order of appearance)), and 1115 (Figure 67C (sequence numbers 1280 to 1281 in order of appearance)). [Figure 67B] Figures 67A to 67C show the NGS read alignments of modeled translocation targets located 60 bp away from the PAM for targets 1093 (Figure 67A (sequence numbers 1264 to 1274 in order of appearance)), 1101 (Figure 67B (sequence numbers 1275 to 1279 in order of appearance)), and 1115 (Figure 67C (sequence numbers 1280 to 1281 in order of appearance)). [Figure 67C] Figures 67A to 67C show the NGS read alignments of modeled translocation targets located 60 bp away from the PAM for targets 1093 (Figure 67A (sequence numbers 1264 to 1274 in order of appearance)), 1101 (Figure 67B (sequence numbers 1275 to 1279 in order of appearance)), and 1115 (Figure 67C (sequence numbers 1280 to 1281 in order of appearance)). [Figure 68] Figure 68 shows NGS detection of mouse albumin (mALB) transposition in Hepa1-6. Guide targets 1141-1182 were predicted sites for mouse albumin intron 1. [Figure 69A]Figures 69A-69C show the NGS read alignments for targets 1148 (Figure 69A (sequence numbers 1282-1285 in order of appearance)), 1161 (Figure 69B (sequence numbers 1286-1288 in order of appearance)), and 1162 (Figure 69C (sequence numbers 1289-1297 in order of appearance)) at a position 60 base pairs away from the PAM. [Figure 69B] Figures 69A-69C show the NGS read alignments for targets 1148 (Figure 69A (sequence numbers 1282-1285 in order of appearance)), 1161 (Figure 69B (sequence numbers 1286-1288 in order of appearance)), and 1162 (Figure 69C (sequence numbers 1289-1297 in order of appearance)) at a position 60 base pairs away from the PAM. [Figure 69C] Figures 69A-69C show the NGS read alignments for targets 1148 (Figure 69A (sequence numbers 1282-1285 in order of appearance)), 1161 (Figure 69B (sequence numbers 1286-1288 in order of appearance)), and 1162 (Figure 69C (sequence numbers 1289-1297 in order of appearance)) at a position 60 base pairs away from the PAM. [Figure 70] Figure 70 shows NGS detection of translocations in mouse albumin (mROSA26) in Hepa1-6. Guide targets 1183-1167 were the predicted sites of mRosa26. [Figure 71A] Figures 71A–71E show the NGS read alignments of modeled dislocation targets located 60 bp away from the target PAM (Figure 71A (Sequence IDs 1298–1304 in visual order)), 1201 (Figure 71B (Sequence IDs 1305–1313 in visual order)), 1205 (Figure 71C (Sequence IDs 1314–1323 in visual order)), 1219 (Figure 71D (Sequence IDs 1324–1327 in visual order)), and 1257 (Figure 71E (Sequence IDs 1328–1330 in visual order)). [Figure 71B]Figures 71A–71E show the NGS read alignments of modeled dislocation targets located 60 bp away from the target PAM (Figure 71A (Sequence IDs 1298–1304 in visual order)), 1201 (Figure 71B (Sequence IDs 1305–1313 in visual order)), 1205 (Figure 71C (Sequence IDs 1314–1323 in visual order)), 1219 (Figure 71D (Sequence IDs 1324–1327 in visual order)), and 1257 (Figure 71E (Sequence IDs 1328–1330 in visual order)). [Figure 71C] Figures 71A–71E show the NGS read alignments of modeled dislocation targets located 60 bp away from the target PAM (Figure 71A (Sequence IDs 1298–1304 in visual order)), 1201 (Figure 71B (Sequence IDs 1305–1313 in visual order)), 1205 (Figure 71C (Sequence IDs 1314–1323 in visual order)), 1219 (Figure 71D (Sequence IDs 1324–1327 in visual order)), and 1257 (Figure 71E (Sequence IDs 1328–1330 in visual order)). [Figure 71D] Figures 71A–71E show the NGS read alignments of modeled dislocation targets located 60 bp away from the target PAM (Figure 71A (Sequence IDs 1298–1304 in visual order)), 1201 (Figure 71B (Sequence IDs 1305–1313 in visual order)), 1205 (Figure 71C (Sequence IDs 1314–1323 in visual order)), 1219 (Figure 71D (Sequence IDs 1324–1327 in visual order)), and 1257 (Figure 71E (Sequence IDs 1328–1330 in visual order)). [Figure 71E] Figures 71A–71E show the NGS read alignments of modeled dislocation targets located 60 bp away from the target PAM (Figure 71A (Sequence IDs 1298–1304 in visual order)), 1201 (Figure 71B (Sequence IDs 1305–1313 in visual order)), 1205 (Figure 71C (Sequence IDs 1314–1323 in visual order)), 1219 (Figure 71D (Sequence IDs 1324–1327 in visual order)), and 1257 (Figure 71E (Sequence IDs 1328–1330 in visual order)). [Figure 72]Figure 72 shows the NGS reads of the conditions as the number of copies of the single guide increases. The control condition reflects pHelper having all components expressed on the plasmid (a single guide targeting Cas12k, S15, TniQ, TnsB, TnsC, and AAVS1 5) using a donor plasmid containing a fluorescent marker. For the remaining conditions, FIX cargo was used to quantify integration. The 0×sg condition did not have a single guide on the pHelper plasmid using FIX pDonor. The 1×sg condition had a single guide on the pHelper plasmid with pDonor FIX. The 2×sg condition did not contain a single guide on the pHelper, which had two single guide cassettes targeting AAVS1-5 on the pDonor FIX plasmid. The 3×sg condition contained pDonor FIX containing two single guide cassettes for AAVS1-5, and the pHelper plasmid contained an additional single guide cassette targeting AAVS1 5. [Figure 73] Figure 73 shows the ddPCR detection results of the LE and RE junctions in promoter-driven FIX (factor IX) delivery, comparing the results under non-induction conditions (pHelper Null) and AAVS1-induction conditions (pHelper AAVS5). [Figure 74] Figure 74 shows that MG161 family members are distant homologs of sso7d. The lineage was inferred from multiple sequence alignments of full-length protein sequences containing PFam PF02294 domain hits. The reference sso7d sequence is highlighted with triangles. The distance between tips is estimated to be 0.5 substitutions per site (horizontal bars). [Figure 75A] Figures 75A–75C show the MG161 functional domain encoded as a tandem repeat. Figure 75A shows the genomic environment of a protein encoding multiple functional domains (FDs). The FDs correspond to tandem incomplete repeats (arrows labeled 161-12 to 161-18). [Figure 75B]Figures 75A–75C show the MG161 functional domain encoded as a tandem repeat. Figure 75B shows multiple sequence alignments of the tandem repeat FD relative to the reference sso7d sequence from S. solfataricus (sequence numbers 1331–1339, respectively, in visual order). MG161–13 has 20% AAI relative to the reference sequence, while the other FDs have lower sequence homology. [Figure 75C] Figures 75A–75C show the MG161 functional domains coded as tandem repeats. Figure 75C shows the 3D structural prediction of the ORF coding the repeating domains, showing that each domain forms a 4-beta sheet structure linked by flexible linkers. [Figure 76] Figure 76 shows that MG162 family members are distant homologs of HMGN1. The lineage was inferred from a multiplex sequence alignment of full-length protein sequences containing the PFam PF01101 domain (sequence numbers 1340-1353 in order of appearance). The reference HMGN1 sequence is highlighted with a triangle. The distance between tips is estimated to be 0.7 substitutions per site (horizontal bars). [Figure 77-1] Figure 77 shows multiple sequence alignments of the MG162 functional domain protein pair with a reference human sequence and a mouse HMGN1 sequence. The mean pairwise percentage identity of the alignments is 40.4%. The conserved RXSXLRS motif (SEQ ID NO: 1255) is highlighted with a black box. [Figure 77-2] Figure 77 shows multiple sequence alignments of the MG162 functional domain protein pair with a reference human sequence and a mouse HMGN1 sequence. The mean pairwise percentage identity of the alignments is 40.4%. The conserved RXSXLRS motif (SEQ ID NO: 1255) is highlighted with a black box. [Figure 78A]Figures 78A–78B show the TniQ assay of fused functional domains. Figure 78A shows a library of functional domains introduced into a TniQ expression vector and separately expressed in a pHelper vector (a single guide targeting Cas12k, S15, TnsB, TnsC, and AAVS1–5) containing other components necessary for transduction. The wild-type H1 core TniQ fusion and pHelper split vector controls were undetectable. Figure 78B shows the top 14 most active functional domains from the split vector assay retested with single-expression constructs. Of all the constructs tested, 13 of the 14 functional domains in the single-expression context demonstrated greater translocation efficiency than the WT single-expression control. [Figure 78B] Figures 78A–78B show the TniQ assay of fused functional domains. Figure 78A shows a library of functional domains introduced into a TniQ expression vector and separately expressed in a pHelper vector (a single guide targeting Cas12k, S15, TnsB, TnsC, and AAVS1–5) containing other components necessary for transduction. The wild-type H1 core TniQ fusion and pHelper split vector controls were undetectable. Figure 78B shows the top 14 most active functional domains from the split vector assay retested with single-expression constructs. Of all the constructs tested, 13 of the 14 functional domains in the single-expression context demonstrated greater translocation efficiency than the WT single-expression control.
[0049] A brief explanation of sequence listings The sequence listings submitted herein provide exemplary polynucleotide and polypeptide sequences for use in the methods, compositions, and systems described herein. The following is an exemplary description of some of these sequences.
[0050] MG64 Sequence IDs 1, 12, 16, 20-30, 64, 80-85, and 220 show the full-length peptide sequences of the MG64 Cas effector.
[0051] Sequence IDs 2-4, 13-15, 17-19, 65-67, and 109-111 show peptide sequences of MG64 transposition proteins that may contain a recombinase / transposase recognition complex associated with the MG64 Cas effector.
[0052] Sequence IDs 5-6, 32-33, 94-95, 104-105, 119-122, and 222 show the nucleotide sequences of MG64 tracrRNA derived from the same locus as the MG64 Cas effector.
[0053] Sequence IDs 7 and 34-35 show the nucleotide sequences of the MG64-targeted CRISPR repeat.
[0054] Sequence IDs 106-108, 112-118, and 221 show the nucleotide sequences of MG64 crRNA.
[0055] Sequence IDs 8, 10, 39-44, 77, 79, and 93 show the nucleotide sequences of right-side transposase recognition sequences associated with the MG64 system.
[0056] Sequence IDs 9, 11, 36-38, 76, and 78 show the nucleotide sequences of the left-side transposase recognition sequences associated with the MG64 system.
[0057] Sequence IDs 45–63, 68–75, 96–103, and 123–140 show the nucleotide sequences of single guide RNAs engineered to function with the MG64 Cas effector.
[0058] Sequence ID 208 shows the nucleotide sequence of the MG64 expression construct.
[0059] Sequence ID 223 shows the nucleotide sequence of the MG64 active donor.
[0060] Sequence IDs 228-230 show the full-length peptide sequences of the MG64 accessory protein.
[0061] Sequence IDs 233-234 show the nucleotide sequences of the MG64 target site.
[0062] Sequence IDs 369-371 show the nucleotide sequences of the MG64 active target sequence.
[0063] MG190 Sequence IDs 209-219 show the full-length peptide sequences of the MG190 ribosomal protein S15 homolog.
[0064] Other arrays Sequence IDs 86-87, 192-207, and 1354-1383 represent the peptide sequences of the nuclear localization signal.
[0065] Sequence IDs 88-89 show the linker peptide sequences.
[0066] Sequence IDs 90-92 show the peptide sequences of the epitope tag.
[0067] Sequence numbers 141-143 represent genome target sequences.
[0068] Sequence numbers 144-180 represent the target guide sequences.
[0069] Sequence IDs 181-183 show the nucleic acid sequences of the S15 fusion protein.
[0070] Sequence ID 184 represents the donor construct.
[0071] Sequence ID 185 shows the MG64-1 sgRNA sequence.
[0072] Sequence ID 186 shows the linker sequence.
[0073] Sequence IDs 187-189 show the amino acid sequences of the S15 fusion protein.
[0074] Sequence numbers 190-191 represent promoter sequences.
[0075] Sequence IDs 224-226 and 231-232 show the nucleotide sequences of the primers.
[0076] Sequence ID 227 shows the nucleotide sequence of the plasmid element.
[0077] Sequence IDs 235-249 show the peptide sequences of the ClpX accessory protein.
[0078] Sequence IDs 250-251 show the nucleotide sequences of the primers.
[0079] Sequence ID 252 shows the nucleotide sequence of the plasmid element.
[0080] Sequence ID 253 shows the nucleotide sequence of the primer.
[0081] Sequence IDs 254-256 show the nucleotide sequences of the MG64 active target sequence.
[0082] Sequence IDs 257-282 and 1138-1241 show the protein sequences of the MG161 functional domain.
[0083] Sequence IDs 283-307 and 1242-1254 show the protein sequences of the MG162 functional domain.
[0084] Sequence IDs 308-359 show the protein sequences of the H1 core library.
[0085] Sequence IDs 360-368 and 372-753 show the nucleotide sequences of the primers and probes.
[0086] Sequence IDs 754-944 represent the nucleotide sequences of a single guide target.
[0087] Sequence IDs 945-1135 show the nucleotide sequences of the primer-binding sequences.
[0088] Sequence ID 1136 shows the nucleotide sequence of the promoter.
[0089] Sequence ID 1137 shows the protein sequence of the expression construct. [Modes for carrying out the invention]
[0090] While various embodiments of the Disclosure are shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided only as examples. Numerous variations, modifications, and substitutions can be conceived by those skilled in the art without departing from the Disclosure. It should be understood that various alternatives to the embodiments of the Disclosure described herein may be used.
[0091] The practices of some of the methods disclosed herein, unless otherwise indicated, utilize techniques from immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA. See, for example, Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (FMAusubel, et al. eds.); the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (MJ MacPherson, BD Hames and GRTaylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (RIFreshney, ed. (2010)).
[0092] As used herein, the singular forms “a,” “an,” and “the” (“one”) are intended to include the plural form unless the context otherwise explicitly indicates otherwise. Furthermore, the terms “including,” “includes,” “having,” “has,” and “with,” or their variations thereof, are intended to be as inclusive as the term “comprising,” to the extent that they are used in either the detailed description and / or the claims.
[0093] The terms “about” or “approximately” mean within an acceptable range of error for a particular value, as determined by those skilled in the art, which depends in part on how the value is measured or determined, i.e., on the limits of the measuring system. For example, “about” may mean within one or more standard deviations according to the practice of the art. Alternatively, “about” may mean a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of a given value.
[0094] As used in this disclosure, the term “nucleotide” refers to a base-sugar-phosphate combination. Nucleotides intended to be nucleotides include naturally occurring and synthetic nucleotides. A nucleotide is the monomeric unit of a nucleic acid sequence (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide includes ribonucleoside triphosphates such as adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP), and deoxyribonucleoside triphosphates, e.g., dATP, dCTP, dITP, dUTP, dGTP, dTTP, or their derivatives. Such derivatives include, for example, [αS]dATP, 7-deaza-dGTP, and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. As used in this disclosure, the term nucleotide also includes dideoxyribonucleoside triphosphate (ddNTP) and its derivatives. Exemplary examples of ddNTPs include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides may be unlabeled or may be labeled in a detectable manner, such as by using optically detectable moieties (e.g., fluorophores) or moieties containing quantum dots. Examples of detectable labels include radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzymatic labels. Examples of fluorescent labels for nucleotides include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4'-dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP, available from Perkin Elmer (Foster City, Calif), FluoroLink DeoxyNucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP, fluorescein-15-dATP, fluorescein-12-dUTP, tetramethylrhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2'-dATP, available from Boehringer (Mannheim, Indianapolis, India), and Molecular Chromosome-labeled nucleotides available from Probes (Eugene, Oreg) include BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP. The term nucleotide encompasses chemically modified nucleotides. An exemplary chemically modified nucleotide is biotin-dNTP.Non-limiting examples of biotinylated dNTPs include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).
[0095] The terms “polynucleotide,” “oligonucleotide,” and “nucleic acid” are used interchangeably to refer to polymeric forms of nucleotides of any length, which are either deoxyribonucleotides or ribonucleotides, or analogues thereof, in single-stranded, double-stranded, or multi-stranded forms. The polynucleotides intended include genes or fragments thereof. Exemplary polynucleotides include, but are not limited to, DNA, RNA, coding or non-coding regions of genes or gene fragments, multiple gene loci (one locus) defined from binding analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, cell-free polynucleotides including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. When referring to T, in polynucleotides, T means U (uracil) in RNA and T (thymine) in DNA. Polynucleotides can be exogenous or endogenous to cells and / or present in a cell-free environment. The term polynucleotide encompasses modified polynucleotides (e.g., modified backbones, sugars, or nucleic acid bases). Where present, modifications to the nucleotide structure are conferred before or after polymer assembly. Non-limiting examples of modifications include 5-bromouracil, peptide nucleic acids, heteronucleotides, morpholino, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein conjugated to sugars), thiol-containing nucleotides, biotin-conjugated nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queosin, and waiosin. The sequence of nucleotides can be interrupted by non-nucleotide components.
[0096] The terms “peptide,” “polypeptide,” and “protein” are used interchangeably herein to refer to polymers of at least two amino acid residues joined by peptide bonds. These terms do not imply a specific length of the polymer and are not intended to imply or distinguish whether peptides are produced using recombinant techniques, chemical or enzymatic synthesis, or naturally occurring. The terms apply to naturally occurring amino acid polymers as well as amino acid polymers containing at least one modified amino acid. In some cases, the polymer is interrupted by non-amino acids. The terms include amino acid chains of any length, including full-length proteins and proteins (e.g., domains) with or without secondary or tertiary structures. The terms also encompass amino acid polymers modified by any other operations, such as disulfide bond formation, glycosylation, lipid formation, acetylation, phosphorylation, oxidation, and conjugation with labeling components. As used in this disclosure, the terms “amino acid” and “multiple amino acids” refer to natural and non-natural amino acids, including but not limited to modified amino acids. Modified amino acids include amino acids that have been chemically modified to include a group or chemical moiety that is not naturally present on the amino acid. The term "amino acid" includes both D-amino acids and L-amino acids.
[0097] As used in this disclosure, “non-natural” means a nucleic acid or polypeptide sequence that does not exist in nature. Non-natural means a nucleic acid or polypeptide sequence that does not exist in nature, including modifications such as mutations, insertions, or deletions. The term non-natural encompasses fusion nucleic acids or polypeptides in which the non-natural sequence encodes or exhibits activity (e.g., enzyme activity, methyltransferase activity, acetyltransferase activity, kinase activity, ubiquitination activity, etc.) of the nucleic acid or polypeptide sequence to which it is fused. Non-natural nucleic acids or polypeptide sequences include those that are genetically engineered to ligate a naturally occurring nucleic acid or polypeptide sequence (or a variant thereof) to produce a chimeric nucleic acid or polypeptide sequence encoding a chimeric nucleic acid or polypeptide.
[0098] As used herein, “operably linked,” “operable linkage,” “operatively linked,” or their grammatical equivalents refer to the arrangement of gene elements, such as promoters, enhancers, polyadenylation sequences, etc., where the action of a first gene element (e.g., movement or activation) has some effect on a second gene element. The effect on the second gene element may, but does not have to be, the same type as the action of the first gene element. For example, if the movement of the first element causes the activation of the second element, then the two gene elements are operably linked. For example, if a regulatory element, which may include a promoter sequence and / or an enhancer sequence, helps initiate transcription of a coding sequence, then the regulatory element is operably linked to the coding region. Intervening residues may exist between the regulatory element and the coding region as long as this functional relationship is maintained.
[0099] A “functional fragment” of a DNA or protein sequence refers to a fragment that possesses biological activity (either functional or structural) substantially similar to that of the full-length DNA or protein sequence. The biological activity of a DNA sequence includes its ability to influence expression in a manner attributable to the full-length sequence.
[0100] The terms “engineered,” “synthetic,” and “artificial” are used interchangeably herein to refer to objects modified by human intervention. For example, these terms refer to polynucleotides or polypeptides that do not exist in nature. Engineered peptides have, but do not require, low sequence homology to naturally occurring human proteins (e.g., less than 50%, less than 25%, less than 10%, less than 5%, less than 1%). For example, the VPR domain and VP64 domain are synthetic transactivation domains. Non-limiting examples include: nucleic acids modified by altering their sequence to a sequence that does not occur in nature; nucleic acids modified by ligating them to nucleic acids that are not related in nature so that the ligated product possesses a function not present in the original nucleic acid; engineered nucleic acids synthesized in vitro using sequences that do not exist in nature; proteins modified by altering their amino acid sequence to a sequence that does not exist in nature; engineered proteins that acquire new functions or properties. An “engineered” system includes at least one engineered component.
[0101] The terms "tracrRNA" or "tracr sequence" refer to the transactivation of CRISPR RNA. TracrRNA interacts with CRISPR(cr)RNA to form the guide (g)RNA for type II and subtype VB CRISPR-Cas systems. When tracrRNA is manipulated, it may have approximately 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% sequence homology and / or similarity to a wild-type exemplary tracrRNA sequence (e.g., tracrRNA derived from S. pyogenes, S. aureus). TracrRNA may refer to modified forms of tracrRNA, which may include nucleotide changes such as deletions, insertions, or substitutions, variants, mutations, or chimeric forms. The term tracrRNA encompasses nucleic acids that may be at least approximately 60% identical to a wild-type exemplary tracrRNA sequence (e.g., tracrRNA from S. pyogenes, S. aureus, etc.) over a stretch of at least six consecutive nucleotides. For example, a tracrRNA sequence may have at least approximately 60% identity, at least approximately 65% identity, at least approximately 70% identity, at least approximately 75% identity, at least approximately 80% identity, at least approximately 85% identity, at least approximately 90% identity, at least approximately 95% identity, at least approximately 98% identity, at least approximately 99% identity, or 100% identity over a stretch of at least six consecutive nucleotides to a wild-type exemplary tracrRNA sequence (e.g., tracrRNA from S. pyogenes, S. aureus, etc.). Type II tracrRNA sequences can be predicted on a genomic sequence by identifying regions that are complementary to a portion of the repetitive sequences in an adjacent CRISPR array.
[0102] As used herein, “guide nucleic acid” or “guide polynucleotide” refers to a nucleic acid that hybridizes to a target nucleic acid, thereby directing the associated nuclease to the target nucleic acid. A guide nucleic acid is, but is not limited to, RNA (guide RNA or gRNA), DNA, or a mixture of RNA and DNA. A guide nucleic acid may include crRNA or tracrRNA, or a combination of both. The term guide nucleic acid encompasses engineered guide nucleic acids and programmable guide nucleic acids that specifically bind to the target nucleic acid. A portion of the target nucleic acid may be complementary to a portion of the guide nucleic acid. A double-stranded target polynucleotide chain that is complementary to the guide nucleic acid and hybridizes with it is called the complementary chain. A double-stranded target polynucleotide chain that is complementary to the complementary chain and therefore not complementary to the guide nucleic acid is called the non-complementary chain. A guide nucleic acid having a polynucleotide chain is called a “single guide nucleic acid.” A guide nucleic acid having two polynucleotide chains is called a “double guide nucleic acid.” Unless otherwise specified, the term “guide nucleic acid” is inclusive and refers to both single guide nucleic acids and double guide nucleic acids. A guide nucleic acid may include a segment referred to as a “nucleic acid targeting segment,” “nucleic acid targeting sequence,” or “spacer.” A nucleic acid targeting segment may include a subsegment referred to as a “protein-binding segment,” “protein-binding sequence,” or “Cas protein-binding segment.”
[0103] The term “sequence homology” or “identity rate” in relation to two or more nucleic acid or polypeptide sequences refers to two or more sequences (e.g., in pairwise alignment) or (e.g., in multiple sequence alignment) that are identical or have a specific proportion of identical amino acid residues or nucleotides when compared and aligned for maximum correspondence across a local or global comparison window, as measured using a sequence comparison algorithm. Suitable sequence comparison algorithms for polypeptide sequences include, for example, BLASTP using a BLOSUM62 scoring matrix with a word length (W) parameter of -3, an expected value (E) parameter of -10, and gap costs set by the presence of -11 and extension of -1, and with conditional composition score matrix adjustment for polypeptide sequences longer than 30 residues; BLASTP using a PAM30 scoring set gap cost with a word length (W) parameter of 2, an expected value (E) parameter of 1,000,000, and open gaps of -9 and extension gaps of -1 for sequences shorter than 30 residues (these are the default parameters for BLASTP in the BLAST suite available at https: / / blast.ncbi.nlm.nih.gov); CLUSTALW using Smith-Waterman homology search algorithm parameters of 2 match, -1 mismatch, and -1 gap; MUSCLE using default parameters; MAFFT using 2 retrieval and 1,000 maximum repeat parameters; Novafold using default parameters; and HMMER hmmalign using default parameters.
[0104] Any variant of the enzymes described herein having one or more conserved amino acid substitutions is included in this disclosure. Such conserved substitutions can be made in the amino acid sequence of a polypeptide without disrupting the three-dimensional structure or function of the polypeptide. Conservative substitutions can be achieved by substituting amino acids that have similar hydrophobicity, polarity, and R-chain length with each other. In addition, or alternatively, conserved substitutions can be identified by comparing aligned sequences of homologous proteins from different species, thereby identifying amino acid residues that are not mutated between species (e.g., residues that are not conserved without altering the fundamental function of the encoded protein). Such conservatively substituted variants may include variants having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of the systems described herein (e.g., the MG64 system described herein). In some embodiments, such conservatively substituted variants are functional variants. Such functional variants may include sequences with substitutions such that the activity of key active site residues of the endonuclease is not disrupted. In some embodiments, any functional variant of the systems described herein lacks at least one substitution of the conserved or functional residues called out in Figures 4A, 4B, and 5. In some embodiments, any functional variant of the systems described herein lacks all substitutions of the conserved or functional residues called out in Figures 4A, 4B, and 5.
[0105] The disclosure also includes variants of any of the enzymes described herein (e.g., deactivation variants) having substitutions of one or more catalytic residues to reduce or eliminate the activity of the enzyme. In some embodiments, the deactivation variant as a protein described herein includes at least one, at least two, or all three catalytic residues.
[0106] Conservative substitution tables that provide functionally similar amino acids are available from various references (e.g., Creighton, Proteins: Structures and Molecular Properties (WH Freeman & Co.); 2 nd See Edition (December 1993). The following eight groups each contain amino acids that are conserved substitutions with each other: 1) Alanine (A), Glycine (G); 2) Aspartic acid (D), glutamic acid (E); 3) Asparagine (N), glutamine (Q); 4) Arginine (R), Lysine (K); 5) Isoleucine (I), leucine (L), methionine (M), valine (V); 6) Phenylalanine (F), tyrosine (Y), tryptophan (W); 7) Serine (S), threonine (T); and 8) Cysteine (C), Methionine (M).
[0107] As used herein, the term “RuvC_III domain” refers to the third discontinuous segment of the RuvC endonuclease domain (the RuvC nuclease domain consists of three discontinuous segments: RuvC_I, RuvC_II, and RuvC_III). The RuvC domain or its segments can generally be identified by alignment to a documented domain sequence, structural alignment to a protein with an annotated domain, or comparison with a hidden Markov model (HMM) constructed based on a documented domain sequence (e.g., Pfam HMM PF18541 of RuvC_III).
[0108] As used in this disclosure, the term “HNH domain” refers to an endonuclease domain having characteristic histidine and asparagine residues. HNH domains can generally be identified by alignment to a documented domain sequence, structural alignment to a protein having an annotated domain, or comparison to a hidden Markov model (HMM) constructed based on a documented domain sequence (e.g., Pfam HMM PF01844 of the HNH domain).
[0109] As used herein, the term “recombinase” refers to an enzyme that mediates the recombination of DNA fragments located between recombinase recognition sequences, resulting in the excision, insertion, inversion, exchange, or transposition of DNA fragments located between recombinase recognition sequences.
[0110] As used herein, the terms “recombinate” or “recombinant” in relation to nucleic acid modification (e.g., genome modification) refer to the process by which two or more nucleic acid molecules or two or more regions of a single nucleic acid molecule are modified by the action of a recombinase protein. Recombination can, in particular, result in, for example, the excision, insertion, inversion, exchange, or transposition of nucleic acid sequences in or between one or more nucleic acid molecules.
[0111] As used herein, the terms “transposon” or “transposable element” refer to nucleic acid sequences in the genome that are mobile gene elements capable of changing their position within the genome. In some embodiments, transposons transport additional “cargo DNA” excised from the genome. Transposons include, for example, retrotransposons, DNA transposons, autonomous and non-autonomous transposons, and class III transposons. Transposon nucleic acid sequences include, for example, genes encoding congeneral transposases, one or more recognition sequences for transposases, or combinations thereof. In some embodiments, these transposons differ in the type of nucleic acid they transpose, the type of repeats at the ends of the transposon, the type of cargo they carry, or the mode of transposition (i.e., self-repair or host repair). As used herein, the terms “transposase” or “multiple transposases” refer to enzymes that bind to the recognition sequence of a transposon and catalyze its movement to another part of the genome. In some embodiments, the movement is by a cut-and-paste or replication-transposition mechanism.
[0112] As used herein, the terms “Tn7” or “Tn7-like transposase” refer to a family of transposases comprising three main components: a regulatory protein (TnsC) and heteromeric transposases (TnsA and / or TnsB). In addition to the TnsABC transposase protein, the Tn7 element can encode dedicated target site-selective proteins, TnsD and TnsE. Together with TnsABC, the sequence-specific DNA-binding protein TnsD induces transposition into a “Tn7 binding site,” i.e., a conserved site referred to as attTn7. TnsD is a member of a larger family of proteins, which also includes TniQ. TniQ has been shown to target transposition into plasmid degradation sites.
[0113] As used herein, the term “complex” refers to the joining of at least two components. Each of the two components may retain the properties / activities it had before forming the complex. Joining may be by covalent, non-covalent (i.e., hydrogen bonds, ionic interactions, van der Waals interactions, and hydrophobic bonds), the use of a linker, fusion, or any other preferred method. In some embodiments, the components in the complex are polynucleotides, polypeptides, or combinations thereof. For example, the complex may include a Cas protein and a guide nucleic acid.
[0114] In some embodiments, the CAST system described herein comprises one or more Tn7 or Tn7-like transposases. In certain exemplary embodiments, the Tn7 or Tn7-like transposase comprises a multimeric protein complex. In certain exemplary embodiments, the multimeric protein complex comprises TnsA, TnsB, TnsC, or TniQ. In these combinations, the transposases (TnsA, TnsB, TnsC, TniQ) may form complexes or fusion proteins with each other.
[0115] In some embodiments, the CAST system described herein comprises one or more Tn5053 or Tn5053-like transposases. In certain exemplary embodiments, the Tn5053 or Tn5053-like transposase comprises a multimeric protein complex. In certain exemplary embodiments, the multimeric protein complex comprises TnsA, TnsB, TnsC, or TniQ. In these combinations, the transposases (TnsA, TnsB, TnsC, TniQ) may form complexes or fusion proteins with each other.
[0116] As used herein, the term “Cas12k” (or alternatively “Class 2 VK type”) refers to a subtype of the V-type CRISPR system found to be defective in nuclease activity (for example, they may contain at least one defective RuvC domain lacking at least one catalytic residue crucial for DNA cleavage). Such effector subtypes are generally associated with the CAST system.
[0117] In accordance with IUPAC convention, the following abbreviations will be used throughout the examples.
[0118] A = Adenine C = Cytosine G = Guanine T = Chimin R = adenine or guanine Y = Cytosine or Thymine S = Guanine or Cytosine W = Adenine or Thymine K = guanine or thymine M = adenine or cytosine B = C, G, or T D = A, G, or T H = A, C, or T V = A, C, or G
[0119] overview The discovery of novel Cas enzymes with unique functionalities and structures could further disrupt deoxyribonucleic acid (DNA) editing technologies, potentially improving their speed, specificity, functionality, and ease of use. Compared to the predicted prevalence of clustered, regularly spaced short palindromic repeat (CRISPR) systems in microorganisms and the complete diversity of microbial species, the literature has relatively few functionally characterized CRISPR / Cas enzymes. This is partly because, under laboratory conditions, the vast number of microbial species are not readily cultured. Metagenomic sequencing from natural environmental niches representing a large number of microbial species could dramatically increase the number of documented novel CRISPR / Cas systems and potentially accelerate the discovery of new oligonucleotide editing functions. A fruitful recent example of such an approach is demonstrated by the 2016 discovery of the CasX / CasY CRISPR system from metagenomic analysis of natural microbial communities.
[0120] The CRISPR / Cas system is an RNA-directed nuclease complex described as functioning as an adaptive immune system in microorganisms. In their natural context, the CRISPR / Cas system arises in a CRISPR (clustered and regularly arranged short palindromic repeat sequence) operon or locus, which generally consists of two parts: (i) an array of short repeat sequences (30-40 bp) separated by equally short spacer sequences that encode an RNA-based targeting element, and (ii) an ORF encoding a Cas that encodes a nuclease polypeptide directed by the RNA-based targeting element, aligned with an accessory protein / enzyme. Efficient nuclease targeting of a particular target nucleic acid sequence generally requires both (i) complementary hybridization between the first 6-8 nucleic acids of the target (target seed) and a crRNA guide; and (ii) the presence of a protospacer-adjacent motif (PAM) sequence within a defined neighborhood of the target seed (PAMs are typically sequences not commonly represented in the host genome). Depending on the precise function and structure of the system, CRISPR-Cas systems are generally classified into two classes, five types, and 16 subtypes based on shared functional characteristics and evolutionary similarities (see Figure 1).
[0121] Class 1 CRISPR-Cas systems have large, multi-subunit effector complexes and include types I, III, and IV.
[0122] The type I CRISPR-Cas system is considered to have moderate complexity in terms of its components. In the type I CRISPR-Cas system, an array of RNA targeting elements is transcribed as a long precursor crRNA (precrRNA) processed by repeat elements, and a short mature crRNA that orients the nuclease complex to the nucleic acid target is subsequently released, followed by a preferred short consensus sequence called a protospacer-adjacent motif (PAM). This processing occurs via the endoribonuclease subunit (Cas6) of a large endonuclease complex called a cascade, which also includes the nuclease (Cas3) protein component of the crRNA-directed nuclease complex. Cas I nucleases primarily function as DNA nucleases.
[0123] The type III CRISPR system can be characterized by the presence of a central nuclease known as Cas10, along with a repeat-associated mysterious protein (RAMP) containing a Csm or Cmr protein subunit. Similar to the type I system, mature crRNA is processed from precrRNA using a Cas6-like enzyme. Unlike the type I and II systems, the type III system is thought to target and cleave DNA-RNA double strands (such as the DNA strand used as a template for RNA polymerase).
[0124] The type IV CRISPR-Cas system harbors an effector complex containing a highly reduced large subunit nuclease (csf1), two genes for RAMP proteins from the Cas5 (csf3) and Cas7 (csf2) groups, and, in some embodiments, a gene for a predicted smaller subunit. Such systems are generally found on endogenous plasmids.
[0125] Class 2 CRISPR-Cas systems generally possess a single polypeptide multi-domain nuclease effector and include types II, V, and VI.
[0126] The type II CRISPR-Cas system is considered the simplest in terms of its components. In the type II CRISPR-Cas system, processing of mature crRNA with a CRISPR array does not require the presence of a special endonuclease subunit, but rather a small transcoding crRNA (tracrRNA) with a region complementary to the array repeat sequence. The tracrRNA interacts with both its corresponding effector nuclease (e.g., Cas9) and the repeat sequence to form a precursor dsRNA structure, which is cleaved by endogenous RNAse III to produce a mature effector enzyme loaded with both tracrRNA and crRNA. The type II nuclease is known as a DNA nuclease. The type II effector generally exhibits a structure consisting of a RuvC-like endonuclease domain that employs an RNase H fold having an unrelated HNH nuclease domain inserted into the fold of a RuvC-like nuclease domain. The RuvC-like domain is involved in cleaving target (e.g., crRNA-complementary) DNA strands, while the HNH domain is involved in cleaving displaced DNA strands.
[0127] Type V CRISPR-Cas systems are characterized by nuclease effector structures (e.g., Cas12) that are similar to those of type II effectors, containing a RuvC-like domain. Like type II, most (but not all) type V CRISPR systems use tracrRNA to process precrRNA into mature crRNA; however, unlike type II systems which require RNAse III to cleave precrRNA into multiple crRNAs, type V systems are capable of cleaving precrRNA using the effector nuclease itself. Similar to type II CRISPR-Cas systems, type V CRISPR-Cas systems are again known as DNA nucleases. Unlike type II CRISPR-Cas systems, some type V enzymes (e.g., Cas12a) appear to possess robust single-strand nonspecific deoxyribonuclease activity, activated by the first crRNA-directed cleavage of a double-stranded target sequence.
[0128] CRISPR-Cas systems of type VI possess RNA-inducible RNA endonucleases. Instead of a RuvC-like domain, single polypeptide effectors of type VI systems (e.g., Cas13) contain two HEPN ribonuclease domains. Unlike both type II and type V systems, type VI systems also appear to not require tracrRNA to process precrRNA into crRNA. However, similar to type V systems, some type VI systems (e.g., C2C2) appear to possess robust single-strand nonspecific nuclease (ribonuclease) activity, activated by the first crRNA-directed cleavage of the target RNA.
[0129] Due to their simpler structure, Class 2 CRISPR-Cas have been the most widely adopted for manipulation and development as design nucleases / genome editing applications.
[0130] One of the initial adaptations of such a system for in vivo use is (i) recombinantly expressed purified full-length Cas9 (e.g., class 2 type II Cas enzyme) isolated from S. pyogenes SF370, (ii) purified mature crRNA of approximately 42 nt (total crRNA transcribed in vitro from a synthetic DNA template carrying a T7 promoter sequence) carrying a 5' sequence of approximately 20 nt complementary to the target DNA sequence to be cleaved, followed by a 3' tracr binding sequence, (iii) purified tracrRNA transcribed in vitro from a synthetic DNA template carrying a T7 promoter sequence, and (iv) Mg 2+ The system was subsequently improved and engineered, and involved a crRNA of (ii) ligated to the 5' end of (iii) by a linker (e.g., GAAA) to form a single fusion synthetic guide RNA (sgRNA) capable of inducing Cas9 to the target by itself (compare the upper and lower panels in Figure 2).
[0131] Such engineered systems can be adapted for use in mammalian cells by providing DNA vectors encoding (i) an ORF encoding codon-optimized Cas9 (e.g., a class 2 type II Cas enzyme) under a suitable mammalian promoter having a C-terminal nuclear localization sequence (e.g., SV40 NLS) and a suitable polyadenylation signal (e.g., TK pA signal), and (ii) an ORF encoding sgRNA (having a 5' sequence beginning with G, followed by a 20nt complementary targeting nucleic acid sequence, a linker, and a tracrRNA sequence ligated to a 3' tracr binding sequence) under a suitable polymerase III promoter (e.g., U6 promoter).
[0132] Transposons are mobile elements that can move between locations within the genome. Such transposons have evolved to limit the adverse effects they have on the host. Various regulatory mechanisms are used to maintain transpositions at low frequency and sometimes to coordinate transpositions with various cellular processes. Some prokaryotic transposons can also assemble functions that benefit the host or otherwise help maintain the element. Certain transposons may also have evolved mechanisms of strict control over the selection of target sites, the most notable example being the Tn7 family.
[0133] Transposons Tn7 and similar elements serve as reproductive hosts for antibiotic resistance and disease-causing functions in clinical settings, as well as potentially encoding other adaptive functions in their natural environments. While the Tn7 system almost completely avoids integration into critical host genes, for example, it has evolved mechanisms to maximize elemental dispersion by recognizing mobile plasmids and bacteriophages capable of transferring Tn7 between host bacteria.
[0134] Tn7 and Tn7-like elements may possess one pathway that controls the location and timing of their insertion, inducing insertion into a single conserved site within the bacterial genome, and a second pathway that appears to be adapted to maximize targeting to mobile plasmids capable of transporting the elements between bacteria (see Figure 3). The association between Tn7-like transposons and the CRISPR-Cas system suggests that transposons may possess hijacked CRISPR effectors that generate R-loops at target sites, facilitating transposon diffusion via plasmids and phages.
[0135] MG64 series In some embodiments, MG64 systems for rearranging cargo nucleotide sequences within target nucleic acid sites are provided herein. See Figures 4A-4B.
[0136] In certain embodiments, the Specified Description describes a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, comprising: a) a Cas effector complex comprising a class 2 V-type Cas effector, a microprokaryotic ribosomal protein subunit S15, and an engineered guide polynucleotide hybridizing to the target nucleic acid site; b) a Tn7-type transpose complex configured to bind to the Cas effector complex and comprising TnsB, TnsC, and TniQ components and auxiliary proteins; and c) a double-stranded nucleic acid configured to interact with the Tn7-type transpose complex and comprising a cargo nucleotide sequence.
[0137] In some embodiments, the present disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, comprising: a Cas effector complex comprising a class 2 V-type Cas effector, a microprokaryotic ribosomal protein subunit S15, and an engineered guide polynucleotide that hybridizes to a target nucleic acid site; a Tn7-type transpose complex that binds to the Cas effector complex, comprising an auxiliary protein containing TnsB, TnsC, and TniQ components, and a sequence having at least 70% sequence homology with any one of SEQ ID NOs. 228-230 and 235-249; and a double-stranded nucleic acid that interacts with the Tn7-type transposase complex and contains a cargo nucleotide sequence.
[0138] In some embodiments, the present disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, comprising: a Cas effector complex comprising a class 2 V-type Cas effector, a microprokaryotic ribosomal protein subunit S15, and an engineered guide polynucleotide that hybridizes to the target nucleic acid site; a Tn7-type transpotase complex that binds to the Cas effector complex and comprises a functional domain (FD)-TniQ fusion and an auxiliary protein; and a double-stranded nucleic acid that interacts with the Tn7-type transposase complex and comprises a cargo nucleotide sequence.
[0139] In other embodiments, the present disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, comprising: a Cas effector complex comprising a class 2 V-type Cas effector, a microprokaryotic ribosomal protein subunit S15, and an engineered guide polynucleotide that hybridizes to the target nucleic acid site; a Tn7-type transpotase complex bound to the Cas effector complex and comprising a functional domain (FD)-TniQ fusion, wherein the functional domain (FD) comprises a sequence having at least 80% sequence homology with any one of SEQ ID NOs. 257-307 and 1138-1242, and an auxiliary protein; and a double-stranded nucleic acid that interacts with the Tn7-type transposase complex and comprises a cargo nucleotide sequence.
[0140] In some embodiments, the present disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, comprising: a Cas effector complex comprising a class 2 V-type Cas effector, a microprokaryotic ribosomal protein subunit S15, and an engineered guide polynucleotide that hybridizes to the target nucleic acid site; a Tn7-type transposase complex comprising TnsB, TnsC, and TniQ components and an auxiliary protein, wherein the auxiliary protein comprises a sequence having 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity with any one of SEQ ID NOs. 228-230 and 235-249; and a double-stranded nucleic acid that interacts with the Tn7-type transposase complex and comprises the cargo nucleotide sequence.
[0141] In some embodiments, the system comprises a double-stranded nucleic acid containing a cargo nucleotide sequence. In some embodiments, this cargo nucleotide sequence interacts with a Tn7 or Tn5053 transposase complex. In some embodiments, the system comprises a Cas effector complex. In some embodiments, the Cas effector complex comprises a class 2 V Cas effector and an engineered guide polynucleotide configured to hybridize to a target nucleotide sequence. In some embodiments, the system comprises a Tn7 or Tn5053 transposase complex configured to bind to the Cas effector complex, and the Tn7 or Tn5053 transposase complex comprises a TnsB subunit.
[0142] In some embodiments, the cargo nucleotide sequence is adjacent to the left transposase recognition sequence. In some embodiments, the cargo nucleotide sequence is adjacent to the right transposase recognition sequence. In some embodiments, the cargo nucleotide sequence is adjacent to both the left transposase recognition sequence and the right transposase recognition sequence.
[0143] In some embodiments, the target nucleic acid includes a target nucleic acid site. In some embodiments, the target nucleic acid includes a PAM sequence that fits with a Cas effector complex adjacent to the target nucleic acid site. In some embodiments, the PAM sequence is located at 3' of the target nucleic acid site. In some embodiments, the PAM sequence is located at 5' of the target nucleic acid site.
[0144] In some embodiments, the manipulated guide polynucleotide is configured to bind to a class 2 V-type Cas endonuclease. In some embodiments, the class 2 V-type Cas effector is a class 2 VK-type effector. In some embodiments, a Class 2 V-type Cas effector comprises a polypeptide containing a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with sequence numbers 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector comprises a polypeptide containing a sequence having at least about 70% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector comprises a polypeptide containing a sequence having at least about 75% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector comprises a polypeptide containing a sequence having at least about 80% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector comprises a polypeptide containing a sequence having at least about 85% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the Class 2 V-type Cas effector comprises a polypeptide having at least about 90% identity with sequence numbers 1, 12, 16, 20-30, 64, 80-85, and 220.In some embodiments, a Class 2 V-type Cas effector includes a polypeptide containing a sequence having at least about 91% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector includes a polypeptide containing a sequence having at least about 92% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector includes a polypeptide containing a sequence having at least about 93% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector includes a polypeptide containing a sequence having at least about 94% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector includes a polypeptide containing a sequence having at least about 95% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector includes a polypeptide containing a sequence having at least about 96% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector includes a polypeptide containing a sequence having at least about 97% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector includes a polypeptide containing a sequence having at least about 98% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector comprises a polypeptide containing a sequence having at least about 99% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some cases, a Class 2 V-type Cas effector comprises a polypeptide containing a sequence having 100% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220.
[0145] In some embodiments, the TnsB subunit includes a polypeptide having a sequence that is at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to sequence numbers 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide having a sequence that is at least about 70% identical to sequence numbers 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 75% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 80% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 85% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 90% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 91% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 92% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide having sequences that are at least about 93% identical to sequence numbers 2, 13, 17, and 65.In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 94% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 95% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 96% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 97% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 98% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 99% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide having sequences that are 100% identical to sequence numbers 2, 13, 17, and 65.
[0146] In some embodiments, the functional domain (FD)-TniQ fusion contains a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with SEQ ID NOs. In some embodiments, the FD-TniQ fusion includes a polypeptide containing sequences having at least about 70% identity with SEQ ID NOs. 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion includes a polypeptide containing sequences having at least about 75% identity with SEQ ID NOs. 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion includes a polypeptide containing sequences having at least about 80% identity with SEQ ID NOs. 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion includes a polypeptide containing sequences having at least about 85% identity with SEQ ID NOs. 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion includes a polypeptide containing sequences having at least about 90% identity with SEQ ID NOs. 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion comprises a polypeptide containing sequences having at least about 91% identity with SEQ ID NOs. 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion comprises a polypeptide containing sequences having at least about 92% identity with SEQ ID NOs. 257-307 and 1138-1242.In some embodiments, the FD-TniQ fusion includes a polypeptide containing sequences having at least about 93% identity with SEQ ID NOs. 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion includes a polypeptide containing sequences having at least about 94% identity with SEQ ID NOs. 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion includes a polypeptide containing sequences having at least about 95% identity with SEQ ID NOs. 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion includes a polypeptide containing sequences having at least about 96% identity with SEQ ID NOs. 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion includes a polypeptide containing sequences having at least about 97% identity with SEQ ID NOs. 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion comprises a polypeptide containing sequences having at least about 98% identity with SEQ ID NOs. 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion comprises a polypeptide containing sequences having at least about 99% identity with SEQ ID NOs. 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion comprises a polypeptide containing sequences having 100% identity with SEQ ID NOs. 257-307 and 1138-1242.
[0147] In some embodiments, the Tn7 type transposase complex comprises at least one polypeptide containing a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 70% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 75% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 80% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 85% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises a polypeptide containing sequences having at least about 90% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises a polypeptide containing sequences having at least about 91% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 92% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 93% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 94% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 95% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 96% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 97% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 98% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 99% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises a polypeptide containing sequences having 100% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.
[0148] In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with at least one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 70% sequence homology with at least one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 75% sequence homology with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 80% sequence homology with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 85% identity with any one of sequence numbers 3-4, 14-15, 18-19, 66-67, and 109-111.In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 90% identity with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 91% identity with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 92% identity with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 93% identity with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 94% identity with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 95% identity with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 96% identity with any one of sequence numbers 3-4, 14-15, 18-19, 66-67, and 109-111.In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 97% column identity with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 98% identity with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 99% identity with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having 100% identity with any one of sequence numbers 3-4, 14-15, 18-19, 66-67, and 109-111.
[0149] In some embodiments, the Tn7 type transposase complex includes an auxiliary protein. In some embodiments, the auxiliary protein is ClpX.
[0150] In some embodiments, the auxiliary protein comprises a sequence having at least 80% sequence homology with any one of SEQ ID NOs. 228-230 and 235-249. In some cases, the auxiliary protein comprises at least one polypeptide comprising a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs. In some embodiments, the auxiliary protein comprises a polypeptide containing sequences having at least about 70% identity with SEQ ID NOs. 228-230 and 235-249. In some embodiments, the auxiliary protein comprises a polypeptide containing sequences having at least about 75% identity with SEQ ID NOs. 228-230 and 235-249. In some embodiments, the auxiliary protein comprises a polypeptide containing sequences having at least about 80% identity with SEQ ID NOs. 228-230 and 235-249. In some embodiments, the auxiliary protein comprises a polypeptide containing sequences having at least about 85% identity with SEQ ID NOs. 228-230 and 235-249. In some embodiments, the auxiliary protein comprises a polypeptide containing sequences having at least about 90% identity with SEQ ID NOs. 228-230 and 235-249. In some embodiments, the auxiliary protein comprises a polypeptide containing sequences having at least about 91% identity with SEQ ID NOs. 228-230 and 235-249. In some embodiments, the auxiliary protein comprises a polypeptide containing sequences having at least about 92% identity with SEQ ID NOs. 228-230 and 235-249. In some embodiments, the auxiliary protein comprises a polypeptide containing sequences having at least about 93% identity with SEQ ID NOs. 228-230 and 235-249.In some embodiments, the auxiliary protein comprises a polypeptide containing sequences having at least about 94% identity with SEQ ID NOs. 228-230 and 235-249. In some embodiments, the auxiliary protein comprises a polypeptide containing sequences having at least about 95% identity with SEQ ID NOs. 228-230 and 235-249. In some embodiments, the auxiliary protein comprises a polypeptide containing sequences having at least about 96% identity with SEQ ID NOs. 228-230 and 235-249. In some embodiments, the auxiliary protein comprises a polypeptide containing sequences having at least about 97% identity with SEQ ID NOs. 228-230 and 235-249. In some embodiments, the auxiliary protein comprises a polypeptide containing sequences having at least about 98% identity with SEQ ID NOs. 228-230 and 235-249. In some embodiments, the auxiliary protein comprises a polypeptide containing sequences having at least about 99% identity with SEQ ID NOs. 228-230 and 235-249. In some embodiments, the auxiliary protein comprises a polypeptide containing sequences having 100% identity with SEQ ID NOs. 228-230 and 235-249.
[0151] In some embodiments, the systems disclosed herein include at least one manipulated guide polynucleotide, such as a gRNA.
[0152] In some embodiments, the manipulated guide polynucleotide comprises a sequence containing at least 46 to 80 consecutive nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 70% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 75% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 80% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 85% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 90% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222.In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 91% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some cases, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 92% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 93% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 94% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 95% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 96% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 97% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 98% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222.In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides that have at least about 99% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides that have 100% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222.
[0153] In some embodiments, the manipulated guide polynucleotide comprises a sequence containing at least 46 to 80 consecutive nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least about 46 to 80 consecutive nucleotides that are identical to any one of sequence numbers 45-63, 68-75, 96-103, 123-140, and 185. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least about 46 to 80 consecutive nucleotides that have at least about 70% identity with sequence numbers 45-63, 68-75, 96-103, 123-140, and 185. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least about 46 to 80 consecutive nucleotides that have at least about 75% identity with sequence numbers 45-63, 68-75, 96-103, 123-140, and 185. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 80% identity with SEQ ID NOs. 45-63, 68-75, 96-103, 123-140, and 185.In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 90% identity with SEQ ID NOs. 45-63, 68-75, 96-103, 123-140, and 185. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 91% identity with SEQ ID NOs. 45-63, 68-75, 96-103, 123-140, and 185. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 92% identity with SEQ ID NOs. 45-63, 68-75, 96-103, 123-140, and 185. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 93% identity with SEQ ID NOs. 45-63, 68-75, 96-103, 123-140, and 185. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 94% identity with SEQ ID NOs. 45-63, 68-75, 96-103, 123-140, and 185. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 95% identity with SEQ ID NOs. 45-63, 68-75, 96-103, 123-140, and 185. In some embodiments, the manipulated guide polynucleotide is a guide RNA and comprises a sequence containing at least 46 to 80 consecutive nucleotides having at least about 96% identity with SEQ ID NOs. 45-63, 68-75, 96-103, 123-140, and 185.In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 97% identity with SEQ ID NOs. 45-63, 68-75, 96-103, 123-140, and 185. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 98% identity with SEQ ID NOs. 45-63, 68-75, 96-103, 123-140, and 185. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 99% identity with SEQ ID NOs. 45-63, 68-75, 96-103, 123-140, and 185. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence comprising at least about 46 to 80 consecutive nucleotides having 100% identity with SEQ ID NOs. 45–63, 68–75, 96–103, 123–140, and 185.
[0154] In some embodiments, the manipulated guide polynucleotide is a guide RNA and comprises a sequence containing at least 46 to 80 consecutive nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, and at least about 99% identity with any one of SEQ ID NOs. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least about 46 to 80 consecutive nucleotides that are identical to any one of SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least about 46 to 80 consecutive nucleotides that have at least about 70% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 75% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 80% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944.In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 85% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 90% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 91% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 92% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 93% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 94% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944.In some embodiments, the manipulated guide polynucleotide is a guide RNA containing a sequence of at least 46 to 80 consecutive nucleotides having at least about 95% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 97% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 98% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least about 46 to 80 consecutive nucleotides that have at least about 99% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA and includes a sequence containing at least about 46 to 80 consecutive nucleotides that have 100% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944.
[0155] In some embodiments, the manipulated guide polynucleotide is a guide RNA and contains a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, and at least about 99% identity with any one of SEQ ID NOs. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing a sequence identical to any one of SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing a sequence having at least about 70% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing a sequence having at least about 75% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences having at least about 80% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944.In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences having at least about 90% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences having at least about 91% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences having at least about 92% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences having at least about 93% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences having at least about 94% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences having at least about 95% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences having at least about 96% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences having at least about 97% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944.In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences that have at least about 98% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences that have at least about 99% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences that have 100% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944.
[0156] In some embodiments, the present disclosure relates to a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, comprising: a Cas effector complex comprising: i) a class 2 V-type Cas effector having at least 80% sequence homology with any one of sequence numbers 1, 81, 82, 83, and 85; and ii) an engineered guide polynucleotide having at least 80% identity with any one of sequence numbers 5, 6, 45-63, 68-75, 96-103, 123-140, and 754-944; and a Tn7-type transposase complex bound to the Cas effector complex, comprising TnsB, TnsC, and TniQ components, and Tn7-type transposase complex, wherein the TnsB, TnsC, or TniQ component is of sequence numbers 2-4. The present invention provides a system comprising: a Tn7 transposase complex comprising a sequence having at least 80% sequence homology with any one of the above, and an auxiliary protein comprising a sequence having at least 80% sequence homology with any one of sequence numbers 228-230 and 235-249; and a double-stranded nucleic acid comprising, in the order from 5' to 3', i) a left transposase recognition sequence comprising a sequence having at least 80% sequence homology with any one of sequence numbers 9, 11, 36, 37, and 38; ii) a cargo nucleotide sequence; and ii) a right transposase sequence comprising a sequence having at least 80% identity with any one of sequence numbers 8, 39-44, and 93.
[0157] In some embodiments, the present disclosure relates to a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, comprising: a Cas effector complex that hybridizes to the target nucleic acid site, comprising: i) a class 2 V-type Cas effector comprising a sequence having at least 80% sequence homology with SEQ ID NO: 12, and iii) an engineered guide polynucleotide having at least 80% identity with any one of SEQ ID NOs: 32, 102, 104, and 107; and a Tn7-type transposase complex that binds to the Cas effector complex and comprises TnsB, TnsC, and TniQ components, wherein TnsB, TnsC, or Tni The present invention provides a system comprising: a Tn7 transposase complex in which the Q component contains a sequence having at least 80% sequence homology with any one of sequence numbers 13 to 15, and the auxiliary protein contains a sequence having at least 80% sequence homology with any one of sequence numbers 228 to 230 and 235 to 249; and a double-stranded nucleic acid that interacts with the Tn7 transposase complex and contains, in the order from 5' to 3', i) a left-side transposase sequence containing a sequence having at least 80% sequence homology with sequence number 76, ii) the cargo nucleotide sequence, and iii) a right-side transposase sequence containing a sequence having at least 80% identity with sequence number 77.
[0158] In some embodiments, the present disclosure relates to a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, comprising: a Cas effector complex that hybridizes to the target nucleic acid site, comprising: i) a class 2 V-type Cas effector comprising a sequence having at least 80% sequence homology with SEQ ID NO: 16, and ii) an engineered guide polynucleotide having at least 80% identity with any one of SEQ ID NOs: 33, 103, 105, and 108; and a Tn7-type transposase complex that binds to the Cas effector complex and comprises TnsB, TnsC, and TniQ components, wherein the TnsB, TnsC, or TniQ component is sequence The present invention provides a system comprising: a Tn7 transposase complex comprising a sequence having at least 80% sequence homology with any one of sequence numbers 17-19, and an auxiliary protein comprising a sequence having at least 80% sequence homology with any one of sequence numbers 228-230 and 235-249; and a double-stranded nucleic acid that interacts with the Tn7 transposase complex and comprises, in the order from 5' to 3', i) a left transposase recognition sequence comprising a sequence having at least 80% sequence homology with sequence number 78, ii) a cargo nucleotide sequence, ii) another cargo nucleotide sequence, and iii) a right transposase sequence comprising a sequence having at least 80% identity with sequence number 79.
[0159] In some embodiments, the system further comprises a PAM sequence compatible with the Cas effector complex. In some embodiments, the PAM sequence includes sequence number 31.
[0160] In some embodiments, the PAM sequence is located approximately 50 to 70 base pairs from the target nucleic acid site. In some embodiments, the PAM sequence is located at 3' of the target nucleic acid site. In some embodiments, the PAM sequence is located at 5' of the target nucleic acid site.
[0161] In some embodiments, the guide RNA includes, but is not limited to, a spacer sequence that binds to a protospacer sequence (target sequence), crRNA, and optional tracrRNA, comprising a variety of structural elements. In some embodiments, the guide RNA comprises crRNA containing a spacer sequence. In some embodiments, the guide RNA further comprises tracrRNA or modified tracrRNA.
[0162] In some embodiments, the system provided herein includes one or more guide RNAs. In some embodiments, the guide RNA includes a sense sequence. In some embodiments, the guide RNA includes an antisense sequence. In some embodiments, the guide RNA includes a nucleotide sequence other than a region complementary or substantially complementary to the region of the target sequence. For example, crRNA is part of the guide RNA, is considered part of the guide RNA, or is included in the guide RNA, e.g., a crRNA:tracrRNA chimera.
[0163] In some embodiments, the guide RNA comprises synthetic or modified nucleotides. In some embodiments, the guide RNA comprises one or more nucleoside linkers modified from natural phosphodiesters. In some embodiments, the nucleoside linkers of the guide RNA, or the entire sequence of nucleotides therein, are modified. For example, in some embodiments, the nucleoside linkage comprises sulfur (S), such as a phosphorothioate nucleoside linkage.
[0164] In some embodiments, the guide RNA comprises modifications to the ribose sugar or nucleic acid base. In some embodiments, the guide RNA comprises one or more nucleosides containing a modified sugar moiety, where the modified sugar moiety is a modification of the sugar moiety compared to the ribose sugar moiety found in deoxyribose nucleic acids (DNA) and RNA. In some embodiments, the modification is within the ribose ring structure. Exemplary modifications include, but are not limited to, substitution with a hexose ring (HNA), a bicyclic ring having a biradical bridge between the C2 and C4 carbons on the ribose ring (e.g., locked nucleic acid (LNA)), or an unbound ribose ring typically lacking a bond between the C2 and C3 carbons (e.g., UNA). In some embodiments, the sugar-modified nucleoside comprises a bicyclohexose nucleic acid or a tricyclic nucleic acid. In some embodiments, the modified nucleoside comprises a nucleoside in which the sugar moiety is replaced with a non-sugar moiety, e.g., peptide nucleic acid (PNA) or morpholino nucleic acid.
[0165] In some embodiments, the guide RNA contains one or more modified sugars. In some embodiments, the sugar modification includes modifications made by altering substituents on the ribose ring to non-hydrogen groups or 2'-OH groups naturally found in DNA and RNA nucleosides. In some embodiments, substituents are introduced at the 2', 3', 4', or 5' positions, or a combination thereof. In some embodiments, the nucleoside having a modified sugar moiety includes 2'-modified nucleosides, e.g., 2'-substituted nucleosides. In some embodiments, the 2'-sugar-modified nucleoside is a nucleoside having a substituent other than -H or -OH at the 2' position (2'-substituted nucleoside), or includes a 2'-linked biradical and includes 2'-substituted nucleosides and LNA (2'-4' biradical-bridged) nucleosides. Examples of 2'-substituted nucleosides include, but are not limited to, 2'-O-alkyl-RNA, 2'-O-methyl-RNA, 2'-alkoxy-RNA, 2'-O-methoxyethyl-RNA (MOE), 2'-amino-DNA, 2'-fluoro-RNA, and 2'-F-ANA nucleosides. In some embodiments, the modification in the ribose group involves a modification at the 2' position of the ribose group. In some embodiments, the modification at the 2' position of the ribose group is selected from the group consisting of 2'-O-methyl, 2'-fluoro, 2'-deoxy, and 2'-O-(2-methoxyethyl).
[0166] In some embodiments, the guide RNA contains one or more modified sugars. In some embodiments, the guide RNA contains only modified sugars. In certain embodiments, the guide RNA contains about 10%, 25%, 50%, 75%, or more than 90% modified sugars. In some embodiments, the modified sugar is a bicyclic sugar. In some embodiments, the modified sugar contains a 2'-O-methoxyethyl group. In some embodiments, the guide RNA contains both internucleoside linker modifications and nucleoside modifications.
[0167] In some embodiments, the guide RNA includes a sequence complementary to a eukaryotic, fungal, plant, mammalian, or human genome polynucleotide sequence. In some embodiments, the guide RNA includes a sequence complementary to a eukaryotic genome polynucleotide sequence. In some embodiments, the guide RNA includes a sequence complementary to a fungal genome polynucleotide sequence. In some embodiments, the guide RNA includes a sequence complementary to a plant genome polynucleotide sequence. In some embodiments, the guide RNA includes a sequence complementary to a mammalian genome polynucleotide sequence. In some embodiments, the guide RNA includes a sequence complementary to a human genome polynucleotide sequence.
[0168] In some embodiments, the guide RNA is 30 to 250 nucleotides long. In some embodiments, the guide RNA is more than 90 nucleotides long. In some embodiments, the guide RNA is less than 245 nucleotides long. In some embodiments, the guide RNA is 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 220, 240, or more than 240 nucleotides long. In the embodiment where guide RNA is added, the number of guide RNA molecules is approximately 30-40, 30-50, 30-60, 30-70, 30-80, 30-90, 30-100, 30-120, 30-140, 30-160, 30-180, 30-200, 30-220, 30-240, 50-60, 50-70, 50-80, 50-90, 50-100, and 50 The lengths of the nucleotides are approximately 120, 50-140, 50-160, 50-180, 50-200, 50-220, 50-240, 100-120, 100-140, 100-160, 100-180, 100-200, 100-220, 100-240, 160-180, 160-200, 160-220, or 160-240.
[0169] In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 75% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 80% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 85% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 90% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 91% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 92% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 93% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 94% identity with sequence numbers 9, 11, 36-38, 76, and 78.In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 95% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 96% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 97% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 98% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 99% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having 100% identity with sequence numbers 9, 11, 36-38, 76, and 78.
[0170] In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 75% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 80% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 85% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 90% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 91% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 92% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 93% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93.In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 94% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 95% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 96% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 97% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 98% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 99% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having 100% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93.
[0171] In some embodiments, the class 2 V-type Cas effector and Tn7-type transposase complex is encoded by a polynucleotide sequence of less than about 20 kilobases, less than about 15 kilobases, less than about 10 kilobases, or less than about 5 kilobases.
[0172] In some embodiments, the Class 2 V-type effector includes a nuclear localization sequence (NLS). In some embodiments, the NLS is located at the N-terminus of the Class 2 V-type effector. In some embodiments, the NLS is located at the C-terminus of the Class 2 V-type effector. In some embodiments, the NLS is located at both the N-terminus and the C-terminus of the Class 2 V-type effector.
[0173] In some embodiments, the NLS includes a sequence that has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of the sequences 192-207 and 1354-1383. In some embodiments, the NLS includes sequences having at least about 85% identity with sequence numbers 192-207 and 1354-1383. In some embodiments, the NLS includes sequences having at least about 90% identity with sequence numbers 192-207 and 1354-1383. In some embodiments, the NLS includes sequences having at least about 91% identity with sequence numbers 192-207 and 1354-1383. In some embodiments, the NLS includes sequences having at least about 92% identity with sequence numbers 192-207 and 1354-1383. In some embodiments, the NLS includes sequences having at least about 93% identity with sequence numbers 192-207 and 1354-1383. In some embodiments, the NLS includes sequences having at least about 94% identity with sequence numbers 192-207 and 1354-1383. In some embodiments, the NLS includes sequences having at least about 95% identity with sequence numbers 192-207 and 1354-1383. In some embodiments, the NLS includes sequences having at least about 96% identity with sequence numbers 192-207 and 1354-1383. In some embodiments, the NLS includes sequences having at least about 97% identity with sequence numbers 192-207 and 1354-1383.In some embodiments, the NLS includes sequences that have at least about 98% identity with sequence numbers 192-207 and 1354-1383. In some embodiments, the NLS includes sequences that have at least about 99% identity with sequence numbers 192-207 and 1354-1383. In some embodiments, the NLS includes sequences that have 100% identity with sequence numbers 192-207 and 1354-1383.
[0174] [Table 1-1] [Table 1-2] [Table 1-3]
[0175] In some embodiments, the Cas effector complex further comprises the microprokaryotic ribosomal protein subunit S15. In some embodiments, the S15 fusion protein is encoded by a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of sequence numbers 181-183. In some embodiments, S15 is encoded by a sequence having at least about 70% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 75% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 80% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 85% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 90% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 91% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 92% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 93% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 94% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 95% identity with sequence numbers 181-183.In some embodiments, S15 is coded by a sequence having at least about 96% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 97% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 98% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 999% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having 100% identity with sequence numbers 181-183.
[0176] In some embodiments, S15 includes a sequence having at least about 70% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 75% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 80% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 85% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 90% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 91% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 92% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 93% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 94% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 95% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 96% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 97% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 98% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 99% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having 100% identity with sequence numbers 187-189.
[0177] In some embodiments, the Cas effector complex comprises one or more linkers that ligate a class 2 type V effector, a microprokaryotic ribosomal protein subunit S15, a transposase, a gRNA, or a combination thereof. In some embodiments, the linker comprises at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, or 400 amino acids. In some embodiments, the linker comprises at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides. In some embodiments, the linker is coded by the sequence of sequence number 186, or by a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with sequence number 186. In some embodiments, the linker is coded by sequence number 186.
[0178] In some embodiments, the present disclosure provides an engineered nuclease system comprising: an endonuclease comprising a RuvC domain, which is a class 2 VK-type Cas effector derived from an uncultured microorganism and having at least 80% identity with any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220; and an engineered guide polynucleotide comprising a sparser sequence that forms a complex with the endonuclease and hybridizes to a target nucleic acid sequence, wherein the engineered guide polynucleotide comprises a sequence having at least about 80% identity with any one of SEQ ID NOs: 754-944.
[0179] Fusion protein In some embodiments, systems for transposing a cargo nucleotide sequence into a target nucleic acid site containing a fusion protein or nucleic acid encoding a fusion protein are described herein. In some embodiments, the fusion protein or nucleic acid encoding a fusion protein includes a class 2 type V effector, a microprokaryotic ribosomal protein subunit S15, a transposase, a gRNA, or a combination thereof. In some embodiments, the fusion protein includes one or more transposases.
[0180] In some embodiments, the nuclear localization sequence (NLS) is fused to a class 2 V-type effector. In some embodiments, the NLS is fused at the N-terminus of the class 2 V-type effector. In some embodiments, the NLS is fused at the C-terminus of the class 2 V-type effector. In some embodiments, the NLS is fused at both the N-terminus and the C-terminus of the class 2 V-type effector.
[0181] In some embodiments, the NLS includes a sequence that has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of the sequences 192-207 and 1354-1383. In some embodiments, the NLS includes sequences having at least about 85% identity with sequence numbers 192-207 and 1354-1383. In some embodiments, the NLS includes sequences having at least about 90% identity with sequence numbers 192-207 and 1354-1383. In some embodiments, the NLS includes sequences having at least about 91% identity with sequence numbers 192-207 and 1354-1383. In some embodiments, the NLS includes sequences having at least about 92% identity with sequence numbers 192-207 and 1354-1383. In some embodiments, the NLS includes sequences having at least about 93% identity with sequence numbers 192-207 and 1354-1383. In some embodiments, the NLS includes sequences having at least about 94% identity with sequence numbers 192-207 and 1354-1383. In some embodiments, the NLS includes sequences having at least about 95% identity with sequence numbers 192-207 and 1354-1383. In some embodiments, the NLS includes sequences having at least about 96% identity with sequence numbers 192-207 and 1354-1383. In some embodiments, the NLS includes sequences having at least about 97% identity with sequence numbers 192-207 and 1354-1383.In some embodiments, the NLS comprises a sequence having at least about 98% identity with SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 99% identity with SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having 100% identity with SEQ ID NOs: 192-207 and 1354-1383.
[0182] In some embodiments, the fusion protein or nucleic acid encoding the fusion protein comprises a fusion of S15 and a nuclear localization sequence (NLS). In some embodiments, the NLS is fused at the N-terminus of S15. In some embodiments, the NLS is fused at the C-terminus of S15. In some embodiments, the NLS is fused at both the N-terminus and C-terminus of S15.
[0183] In some embodiments, the S15 fusion protein further comprises a cleavable peptide. In some embodiments, the peptide is a 2A peptide.
[0184] In some embodiments, the S15 fusion protein is encoded by a sequence having at least 80% sequence homology with any one of SEQ ID NOs. 181–183. In some embodiments, the S15 fusion protein is encoded by a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs. 181–183. In some embodiments, the S15 fusion protein is encoded by a sequence having at least 80% sequence homology with any one of SEQ ID NOs. 181-183. In some embodiments, the S15 fusion protein is encoded by a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs. 181-183 is encoded by a sequence having at least about 70% identity with SEQ ID NOs. 181-183. In some embodiments, S15 is coded by an array having at least about 75% identity with sequence numbers 181-183.In some embodiments, S15 is coded by a sequence having at least about 80% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 85% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 90% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 91% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 92% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 93% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 94% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 95% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 96% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 97% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 98% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having at least about 99% identity with sequence numbers 181-183. In some embodiments, S15 is coded by a sequence having 100% identity with sequence numbers 181-183.
[0185] In some embodiments, the S15 fusion protein contains a sequence having at least about 70% sequence homology with any one of SEQ ID NOs. 187-189. In some embodiments, the S15 fusion protein has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs. 187-189 contains a sequence having at least about 70% identity with SEQ ID NOs. In some embodiments, S15 includes a sequence having at least about 75% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 80% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 85% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 90% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 91% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 92% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 93% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 94% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 95% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 96% identity with sequence numbers 187-189. In some embodiments, S15 includes a sequence having at least about 97% identity with sequence numbers 187-189.In some embodiments, S15 comprises a sequence having at least about 98% identity with SEQ ID NOs: 187 - 189. In some embodiments, S15 comprises a sequence having at least about 99% identity with SEQ ID NOs: 187 - 189. In some embodiments, S15 comprises a sequence having 100% identity with SEQ ID NOs: 187 - 189.
[0186] In some embodiments, the NLS is fused to a transposase. In some embodiments, the transposase is TnsB, TnsC, or TniQ. In some embodiments, the transposase is TnsB. In some embodiments, the transposase is TnsC. In some embodiments, the transposase is TniQ. In some embodiments, the NLS is fused to the N - terminus of the transposase. In some embodiments, the NLS is fused to the C - terminus of the transposase. In some embodiments, the NLS is fused to both the N - terminus and the C - terminus of the transposase.
[0187] In some embodiments, the fusion protein or the nucleic acid encoding the fusion protein comprises a gRNA (e.g., a dual gRNA or a single gRNA) as described herein.
[0188] In some embodiments, the class 2 type V effector, the small prokaryotic ribosomal protein subunit S15, the transposase, the gRNA, or the fusion protein comprises a tag. In some embodiments, the tag is an affinity tag. In some embodiments, the tag is a polypeptide or a polynucleotide. Exemplary affinity tags include, but are not limited to, His tag, Flag tag, Myc tag, MBP tag, and GST tag.
[0189] In some embodiments, a class 2 type V effector, a microprokaryotic ribosomal protein subunit S15, a transposase, or a fusion protein includes a protease cleavage site. Exemplary protease cleavage sites include, but are not limited to, TEV sites, C3 sites, factor Xa sites, and enterokinase sites.
[0190] cell In certain embodiments, cells comprising the system described herein are described herein.
[0191] In some embodiments, the cells may be eukaryotic cells (e.g., plant cells, animal cells, protist cells, or fungal cells), mammalian cells (Chinese hamster ovary (CHO) cells, baby hamster kidney (BHK), human fetal kidney (HEK), mouse myeloma (NS0), or human retinal cells), immortalized cells (e.g., HeLa cells, COS cells, HEK-293T cells, MDCK cells, 3T3 cells, PC12 cells, Huh7 cells, HepG2 cells, K562 cells, N2a cells, or SY5Y cells), insect cells (e.g., Spodoptera frugiperda cells, Trichoplusia ni cells, Drosophila melanogaster cells, S2 cells, or Heliothis virescens cells), or yeast cells (e.g., Saccharomyces). These include cerevisiae cells, Cryptococcus cells, or Candida cells, plant cells (e.g., parenchymal cells, plaque cells, or plaque-walled cells), fungal cells (e.g., Saccharomyces cerevisiae cells, Cryptococcus cells, or Candida cells), or prokaryotic cells (e.g., E. coli cells, Streptococcus bacterial cells, Streptomyces soil bacterial cells, or archaeal cells). In some embodiments, the cells are eukaryotic cells. In some embodiments, the cells are mammalian cells. In some embodiments, the cells are immortalized cells. In some embodiments, the cells are insect cells. In some embodiments, the cells are yeast cells. In some embodiments, the cells are plant cells. In some embodiments, the cells are fungal cells. In some embodiments, the cells are prokaryotic cells.
[0192] In some embodiments, the cells are A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cells, HT1080, HepG2, Huh7, K562, primary cells, or derivatives thereof.
[0193] Delivery and vector In some embodiments, nucleic acid sequences encoding the MG64 system are disclosed herein, including a Class 2 V-type effector, a microprokaryotic ribosomal protein subunit S15, a transposase, a gRNA, a fusion protein, or a gene editing system.
[0194] In some embodiments, the nucleic acid encoding the MG64 system is DNA, such as linear DNA, plasmid DNA, or minicircle DNA. In some embodiments, the nucleic acid encoding the MG64 system is RNA, such as mRNA.
[0195] In some embodiments, the nucleic acid encoding the MG64 system is delivered by a nucleic acid-based vector. In some embodiments, the nucleic acid-based vector is a plasmid (e.g., a circular DNA molecule that can autonomously replicate inside a cell), a cosmid (e.g., a pWE or sCos vector), an artificial chromosome, a human artificial chromosome (HAC), a yeast artificial chromosome (YAC), a bacterial artificial chromosome (BAC), a P1-derived artificial chromosome (PAC), a phagemid, a phage derivative, a bacmid, or a virus. In some embodiments, the nucleic acid-based vectors are pSF-CMV-NEO-NH2-PPT-3XFLAG, pSF-CMV-NEO-COOH-3XFLAG, pSF-CMV-PURO-NH2-GST-TEV, pSF-OXB20-COOH-TEV-FLAG(R)-6His, pCEP4 pDEST27, pSF-CMV-Ub-KrYFP, pSF-CMV-FMDV-daGFP, pEF1a-mCherry-N1 vector, pEF1a-tdTomato vector, pSF-CMV-FMDV-Hygro, pSF-CMV-PGK-Puro, pMCP-tag(m), pSF-CMV-PURO-NH2-CMYC, pSF-OXB20-BetaGal, pSF-OXB20-Fluc, pSF-OXB20, pSF-Tac, and pRI 101-AN. The selection is made from a list consisting of DNA, pCambia2301, pTYB21, pKLAC2, pAc5.1 / V5-His A, and pDEST8.
[0196] In some embodiments, the nucleic acid-based vector includes a promoter. In some embodiments, the promoter is selected from the group consisting of mini-promoters, inducible promoters, constitutive promoters, and their derivatives. In some embodiments, the promoter is selected from the group consisting of CMV, CBA, EF1a, CAG, PGK, TRE, U6, UAS, T7, Sp6, lac, araBad, trp, Ptac, p5, p19, p40, synapsin, CaMKII, GRK1, and their derivatives. In some embodiments, the promoter is the U6 promoter. In some embodiments, the promoter is the CAG promoter. In some embodiments, the promoter is coded by a sequence that has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of the sequences from sequence numbers 190 to 191.
[0197] In some embodiments, the nucleic acid-based vector is a virus. In some embodiments, the virus is an alphavirus, parvovirus, adenovirus, AAV, baculovirus, dengue virus, lentivirus, herpesvirus, poxvirus, anerovirus, bocavirus, vacciniavirus, or retrovirus. In some embodiments, the virus is an alphavirus. In some embodiments, the virus is a parvovirus. In some embodiments, the virus is an adenovirus. In some embodiments, the virus is an AAV. In some embodiments, the virus is a baculovirus. In some embodiments, the virus is a dengue virus. In some embodiments, the virus is a lentivirus. In some embodiments, the virus is a herpesvirus. In some embodiments, the virus is a poxvirus. In some embodiments, the virus is an anerovirus. In some embodiments, the virus is a bocavirus. In some embodiments, the virus is a vacciniavirus. In some embodiments, the virus is a retrovirus.
[0198] In some embodiments, the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV-rh8, AAV-rh1 0, AAV-rh20, AAV-rh39, AAV-rh74, AAV-rhM4-1, AAV-hu37, AAV-Anc80, AAV-Anc80L65, AAV-7m8, AAV-PHP-B, AAV-PHP-EB, AAV-2.5, AAV-2tYF, A The following are examples of the herpesviruses: AV-3B, AAV-LK03, AAV-HSC1, AAV-HSC2, AAV-HSC3, AAV-HSC4, AAV-HSC5, AAV-HSC6, AAV-HSC7, AAV-HSC8, AAV-HSC9, AAV-HSC10, AAV-HSC11, AAV-HSC12, AAV-HSC13, AAV-HSC14, AAV-HSC15, AAV-TT, AAV-DJ / 8, AAV-Myo, AAV-NP40, AAV-NP59, AAV-NP22, AAV-NP66, AAV-HSC16, or derivatives thereof. In some embodiments, the herpesvirus is HSV type 1, HSV-2, VZV, EBV, CMV, HHV-6, HHV-7, or HHV-8.
[0199] In some embodiments, the virus is AAV1 or a derivative thereof. In some embodiments, the virus is AAV2 or a derivative thereof. In some embodiments, the virus is AAV3 or a derivative thereof. In some embodiments, the virus is AAV4 or a derivative thereof. In some embodiments, the virus is AAV5 or a derivative thereof. In some embodiments, the virus is AAV6 or a derivative thereof. In some embodiments, the virus is AAV7 or a derivative thereof. In some embodiments, the virus is AAV8 or a derivative thereof. In some embodiments, the virus is AAV9 or a derivative thereof. In some embodiments, the virus is AAV10 or a derivative thereof. In some embodiments, the virus is AAV11 or a derivative thereof. In some embodiments, the virus is AAV12 or a derivative thereof. In some embodiments, the virus is AAV13 or a derivative thereof. In some embodiments, the virus is AAV14 or a derivative thereof. In some embodiments, the virus is AAV15 or a derivative thereof. In some embodiments, the virus is AAV16 or a derivative thereof. In some embodiments, the virus is AAV-rh8 or a derivative thereof. In some embodiments, the virus is AAV-rh10 or a derivative thereof. In some embodiments, the virus is AAV-rh20 or a derivative thereof. In some embodiments, the virus is AAV-rh39 or a derivative thereof. In some embodiments, the virus is AAV-rh74 or a derivative thereof. In some embodiments, the virus is AAV-rhM4-1 or a derivative thereof. In some embodiments, the virus is AAV-hu37 or a derivative thereof. In some embodiments, the virus is AAV-Anc80 or a derivative thereof. In some embodiments, the virus is AAV-Anc80L65 or a derivative thereof. In some embodiments, the virus is AAV-7m8 or a derivative thereof. In some embodiments, the virus is AAV-PHP-B or a derivative thereof.In some embodiments, the virus is AAV-PHP-EB or a derivative thereof. In some embodiments, the virus is AAV-2.5 or a derivative thereof. In some embodiments, the virus is AAV-2tYF or a derivative thereof. In some embodiments, the virus is AAV-3B or a derivative thereof. In some embodiments, the virus is AAV-LK03 or a derivative thereof. In some embodiments, the virus is AAV-HSC1 or a derivative thereof. In some embodiments, the virus is AAV-HSC2 or a derivative thereof. In some embodiments, the virus is AAV-HSC3 or a derivative thereof. In some embodiments, the virus is AAV-HSC4 or a derivative thereof. In some embodiments, the virus is AAV-HSC5 or a derivative thereof. In some embodiments, the virus is AAV-HSC6 or a derivative thereof. In some embodiments, the virus is AAV-HSC7 or a derivative thereof. In some embodiments, the virus is AAV-HSC8 or a derivative thereof. In some embodiments, the virus is AAV-HSC9 or a derivative thereof. In some embodiments, the virus is AAV-HSC10 or a derivative thereof. In some embodiments, the virus is AAV-HSC11 or a derivative thereof. In some embodiments, the virus is AAV-HSC12 or a derivative thereof. In some embodiments, the virus is AAV-HSC13 or a derivative thereof. In some embodiments, the virus is AAV-HSC14 or a derivative thereof. In some embodiments, the virus is AAV-HSC15 or a derivative thereof. In some embodiments, the virus is AAV-TT or a derivative thereof. In some embodiments, the virus is AAV-DJ / 8 or a derivative thereof. In some embodiments, the virus is AAV-Myo or a derivative thereof. In some embodiments, the virus is AAV-NP40 or a derivative thereof. In some embodiments, the virus is AAV-NP59 or a derivative thereof. In some embodiments, the virus is AAV-NP22 or a derivative thereof.In some embodiments, the virus is AAV-NP66 or a derivative thereof. In some embodiments, the virus is AAV-HSC16 or a derivative thereof.
[0200] In some embodiments, the virus is HSV-1 or a derivative thereof. In some embodiments, the virus is HSV-2 or a derivative thereof. In some embodiments, the virus is VZV or a derivative thereof. In some embodiments, the virus is EBV or a derivative thereof. In some embodiments, the virus is CMV or a derivative thereof. In some embodiments, the virus is HHV-6 or a derivative thereof. In some embodiments, the virus is HHV-7 or a derivative thereof. In some embodiments, the virus is HHV-8 or a derivative thereof.
[0201] In some embodiments, the nucleic acid encoding the MG64 system is delivered by a non-nucleic acid-based delivery system (e.g., a non-viral delivery system). In some embodiments, the non-viral delivery system is a liposome. In some embodiments, the nucleic acid is lipid-related. Lipid-related nucleic acids are, in some embodiments, encapsulated within the aqueous interior of a liposome, dispersed within the lipid bilayer of a liposome, attached to a liposome via binding molecules related to both the liposome and the nucleic acid, confined within a liposome, complexed with a liposome, dispersed in a lipid-containing solution, mixed with a lipid, combined with a lipid, contained as a suspension in a lipid, contained with a micelle, complexed with a micelle, or otherwise lipid-related. In some embodiments, the nucleic acid is contained in lipid nanoparticles (LNPs).
[0202] In some embodiments, the fusion protein or genome editing system is introduced into cells in any preferred manner, either stably or transiently. In some embodiments, the fusion protein or genome editing system is transfected into cells. In some embodiments, cells are transfected or transfected with a nucleic acid construct encoding the fusion protein or genome editing system. For example, cells are transfected (e.g., with a virus encoding the fusion protein or genome editing system) or transfected with the fusion protein or genome editing system, or with a nucleic acid encoding the translated fusion protein or genome editing system (e.g., with a plasmid encoding the fusion protein or genome editing system). In some embodiments, the transfection is stably or transiently transfected. In some embodiments, cells expressing the fusion protein or genome editing system, or containing the fusion protein or genome editing system, are transfected or transfected with one or more gRNA molecules, for example, when the fusion protein or genome editing system contains a CRISPR nuclease. In some embodiments, plasmids expressing fusion proteins or genome editing systems are introduced into cells by electroporation, transient (e.g., lipofection), stable genome integration (e.g., piggybac), and viral transduction (e.g., lentivirus or AAV), or other methods known to those skilled in the art. In some embodiments, the gene editing system is introduced into cells as one or more polypeptides. In some embodiments, delivery is achieved through the use of RNP complexes. For example, methods for delivering polypeptides and / or RNPs into cells by electroporation or cell compression are known in the art.
[0203] Exemplary methods for nucleic acid delivery include lipofection, nucleofection, electroporation, stable genome integration (e.g., piggybac), microinjection, biolistek, virosomes, liposomes, immunoliposomes, polycationic or lipid nucleic acid conjugates, naked DNA, artificial virions, and drug-enhanced incorporation of DNA. Lipofection is described, for example, in U.S. Patents 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam®, Lipofectin®, and SF Cell Line 4D-Nucleofector X Kit® (Lonza)). Cationic and neutral lipids suitable for efficient receptor-recognition lipofection of polynucleotides include lipids of WO91 / 17424 and WO91 / 16024. In some embodiments, delivery is to cells (e.g., in vitro or ex vivo administration) or target tissue (e.g., in vivo administration). In some embodiments, the nucleic acid is contained in liposomes or nanoparticles that specifically target host cells.
[0204] Additional methods for delivering nucleic acids to cells are known to those skilled in the art. See, for example, US2003 / 0087817.
[0205] In some embodiments, this disclosure provides cells containing the vector or nucleic acid described herein. In some embodiments, the cells express a gene editing system or a part thereof. In some embodiments, the cells are human cells. In some embodiments, the cells are genome-edited ex vivo. In some embodiments, the cells are genome-edited in vivo.
[0206] Methods for transposition The present disclosure provides methods for translocating a cargo nucleotide sequence within a target nucleic acid site. In some embodiments, the method comprises expressing the systems described herein within a cell or introducing the systems described herein into a cell. In some embodiments, the method comprises contacting a cell with the systems described herein.
[0207] In some embodiments, the method comprises contacting a double-stranded nucleic acid comprising the cargo nucleotide sequence with a Cas effector complex comprising a class 2 type V Cas effector and at least one engineered guide polynucleotide configured to hybridize to the target nucleotide sequence. In some embodiments, the method comprises contacting a double-stranded nucleic acid comprising the cargo nucleotide sequence with a Tn7-type transposase complex configured to bind to the Cas effector complex, the Tn7-type transposase complex comprising a TnsB subunit. In some embodiments, the method comprises contacting a double-stranded nucleic acid comprising the cargo nucleotide sequence with a double-stranded target nucleic acid comprising the target nucleic acid site.
[0208] In some embodiments, the cargo nucleotide sequence is adjacent to a left transposase recognition sequence. In some embodiments, the cargo nucleotide sequence is adjacent to a right transposase recognition sequence. In some embodiments, the cargo nucleotide sequence is adjacent to both a left transposase recognition sequence and a right transposase recognition sequence. In some embodiments, the method further comprises a PAM sequence that is compatible with the Cas effector complex adjacent to the target nucleic acid site. In some embodiments, the PAM sequence is located 3' of the target nucleic acid site.
[0209] In some embodiments, the manipulated guide polynucleotide is configured to bind to a class 2 V-type Cas endonuclease. In some embodiments, the class 2 V-type Cas effector comprises a polypeptide containing a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with SEQ ID NOs. In some embodiments, a Class 2 V-type Cas effector comprises a polypeptide containing a sequence having at least about 70% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector comprises a polypeptide containing a sequence having at least about 75% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector comprises a polypeptide containing a sequence having at least about 80% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector comprises a polypeptide containing a sequence having at least about 85% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector comprises a polypeptide containing sequences having at least about 90% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector comprises a polypeptide containing sequences having at least about 91% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220.In some embodiments, a Class 2 V-type Cas effector includes a polypeptide containing a sequence having at least about 92% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector includes a polypeptide containing a sequence having at least about 93% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector includes a polypeptide containing a sequence having at least about 94% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector includes a polypeptide containing a sequence having at least about 95% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector includes a polypeptide containing a sequence having at least about 96% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector includes a polypeptide containing a sequence having at least about 97% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector includes a polypeptide containing a sequence having at least about 98% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector includes a polypeptide containing a sequence having at least about 99% identity with SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, a Class 2 V-type Cas effector comprises a polypeptide having sequences having 100% identity with sequence numbers 1, 12, 16, 20-30, 64, 80-85, and 220.
[0210] In some embodiments, the TnsB subunit includes a polypeptide having a sequence that is at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to sequence numbers 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide having a sequence that is at least about 70% identical to sequence numbers 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 75% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 80% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 85% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 90% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 91% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 92% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide having sequences that are at least about 93% identical to sequence numbers 2, 13, 17, and 65.In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 94% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 95% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 96% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 97% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 98% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component includes a polypeptide containing a sequence having at least about 99% identity with SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide having sequences that are 100% identical to sequence numbers 2, 13, 17, and 65.
[0211] In some embodiments, the Tn7 type transposase complex includes at least one polypeptide (e.g., at least one, two, three, four, five, six, or more polypeptides) containing a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 70% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 75% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 80% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 85% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises a polypeptide containing sequences having at least about 90% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises a polypeptide containing sequences having at least about 91% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 92% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 93% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 94% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 95% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 96% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 97% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 98% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex includes a polypeptide containing sequences having at least about 99% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises a polypeptide containing sequences having 100% identity with SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.
[0212] In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs. In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 70% sequence homology with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 75% sequence homology with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 80% sequence homology with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 85% sequence homology with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.In some embodiments, the Tn7 transposase complex independently comprises at least a first polypeptide and a second polypeptide, each containing a sequence having at least 90% sequence homology with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex independently comprises at least a first polypeptide and a second polypeptide, each containing a sequence having at least 91% sequence homology with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex independently comprises at least a first polypeptide and a second polypeptide, each containing a sequence having at least 92% sequence homology with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 93% sequence homology with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 94% sequence homology with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 95% sequence homology with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 96% sequence homology with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.In some embodiments, the Tn7 transposase complex independently comprises at least a first polypeptide and a second polypeptide, each containing a sequence having at least 97% sequence homology with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex independently comprises at least a first polypeptide and a second polypeptide, each containing a sequence having at least 98% sequence homology with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 transposase complex independently comprises at least a first polypeptide and a second polypeptide, each containing a sequence having at least 99% sequence homology with any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having 100% sequence homology to any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.
[0213] In some embodiments, the manipulated guide polynucleotide comprises a sequence containing at least 46 to 80 consecutive nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 70% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 75% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 80% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 85% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 90% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222.In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 91% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 92% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 93% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 94% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 95% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 96% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 97% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides having at least about 98% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222.In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides that have at least about 99% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222. In some embodiments, the manipulated guide polynucleotide includes a sequence containing at least 46 to 80 consecutive nucleotides that have 100% identity with SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222.
[0214] In some embodiments, the manipulated guide polynucleotide is a guide RNA and contains a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, and at least about 99% identity with any one of SEQ ID NOs. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing a sequence identical to any one of SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing a sequence having at least about 70% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing a sequence having at least about 75% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences having at least about 80% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944.In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences having at least about 90% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences having at least about 91% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences having at least about 92% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences having at least about 93% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences having at least about 94% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences having at least about 95% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences having at least about 96% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences having at least about 97% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944.In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences that have at least about 98% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences that have at least about 99% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944. In some embodiments, the manipulated guide polynucleotide is a guide RNA containing sequences that have 100% identity with SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944.
[0215] In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 75% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 80% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 85% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 90% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 91% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 92% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 93% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 94% identity with sequence numbers 9, 11, 36-38, 76, and 78.In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 95% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 96% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 97% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 98% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having at least about 99% identity with sequence numbers 9, 11, 36-38, 76, and 78. In some embodiments, the left-side transposase recognition sequence includes a sequence having 100% identity with sequence numbers 9, 11, 36-38, 76, and 78.
[0216] In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 75% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 80% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 85% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 90% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 91% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 92% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 93% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93.In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 94% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 95% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 96% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 97% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 98% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having at least about 99% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-side transposase recognition sequence includes a sequence having 100% identity with sequence numbers 8, 10, 39-44, 77, 79, and 93.
[0217] In some embodiments, the class 2 V-type Cas effector and Tn7-type transposase complex is encoded by a polynucleotide sequence of less than about 20 kilobases, less than about 15 kilobases, less than about 10 kilobases, or less than about 5 kilobases.
[0218] use The systems of this disclosure can be used for a variety of applications, such as nucleic acid editing (e.g., gene editing) or binding to nucleic acid molecules (e.g., sequence-specific binding). Such systems can be used, for example, to correct (e.g., remove or replace) genetically inherited mutations that may cause disease in a subject; to inactivate genes to confirm their function in cells; as a diagnostic tool to detect disease-causing genetic elements (e.g., via cleavage of reverse-transcribed viral RNA or amplified DNA sequences encoding disease-causing mutations); as an inactivated enzyme combined with a probe to target and detect specific nucleotide sequences (e.g., sequences encoding antibiotic resistance in bacteria); to inactivate viruses by targeting viral genomes or to prevent them from infecting host cells; to add genes or modify metabolic pathways to manipulate organisms to produce valuable small molecules, macromolecules or secondary metabolites; to establish gene-driven elements for evolutionary selection; and / or as a biosensor to detect cellular perturbations by foreign small molecules and nucleotides.
[0219] kit In some embodiments, the Disclosure provides a kit comprising one or more nucleic acid constructs encoding various components of the fusion protein or genome editing system described herein, for example, a nucleotide sequence encoding a fusion protein or component of the genome editing system capable of modifying a target DNA sequence. In some embodiments, the nucleotide sequence includes a heterologous promoter that drives the expression of the RNA genome editing system component.
[0220] In some embodiments, fusion proteins or gene editing systems comprising a Class 2 V-type effector, a microprokaryotic ribosomal protein subunit S15, a transposase, a gRNA, or any combination thereof, as disclosed herein, are incorporated into pharmaceutical, diagnostic, or research kits to facilitate their use in therapeutic, diagnostic, or research applications. The kit may include one or more containers containing any of the vectors disclosed herein, and instructions for use.
[0221] The kit may be designed to facilitate researchers' use of the methods described herein and may take various forms. Each of the components of the kit may be provided in liquid form (e.g., in solution) or solid form (e.g., dry powder), where applicable. In certain cases, some of the components may be configured to be (e.g., in an active form) or otherwise processable by the addition of a suitable solvent or other kind (e.g., water or cell culture medium), which may or may not be provided with the kit. Where used herein, “instructions” defines the components of the instructions and / or promotional materials and may typically be accompanied by written instructions on or associated with the packaging of the herein. Instructions may also include any oral or electronic instructions provided in any format that clearly indicates to the user that the instructions relate to the kit, such as audiovisual (e.g., videotape, DVD, etc.), the internet, and / or web-based communications. In some embodiments, written instructions may be in the form prescribed by a government agency that regulates the manufacture, use, or sale of a pharmaceutical or biological product, and such instructions may also reflect approval by an agency for manufacture, use, or sale for animal administration.
[0222] Examples The following embodiments are given for the purpose of illustrating various embodiments of the Disclosure and are not intended to limit the Disclosure in any way. These embodiments, along with the methods described herein, represent preferred embodiments at present and are illustrative and are not intended to limit the scope of the Disclosure. Modifications and other uses therein, which are included within the spirit of the Disclosure as defined by the claims, will be recalled by those skilled in the art.
[0223] Example 1 - (General Protocol) PAM sequence identification / confirmation of the system described herein The putative endonuclease was expressed in an E. coli lysate-based expression system. The PAM sequence was determined by sequencing plasmids containing randomly generated potential PAM sequences that could be cleaved by the putative nuclease. In this system, the E coli codon-optimized nucleotide sequence encoding the putative nuclease was transcribed and translated in vitro from a PCR fragment under the control of the T7 promoter. A second PCR fragment having a minimal CRISPR array consisting of a repeat-spacer-repeat sequence, followed by the T7 promoter, was transcribed in the same reaction. Successful expression of the endonuclease sequence and the repeat-spacer-repeat sequence in the E. coli lysate-based expression system, followed by CRISPR array processing, provided an active in vitro CRISPR nuclease complex.
[0224] A library of target plasmids containing a spacer sequence matching the one in the preceding minimal array, where the 8N mixed base (potential PAM sequence) was located, was incubated with the product of the expression system reaction. After 1–3 hours, the reaction was stopped, and the DNA was recovered via a DNA cleanup kit. The adapter sequence was blunt-ended ligated to the DNA containing the cleaved active PAM sequence by the endonuclease, while the uncleaved DNA was inaccessible due to ligation. The DNA segments containing the active PAM sequence were then amplified by PCR using primers specific to the library and the adapter sequence. The PCR amplification products were degraded on a gel to identify the amplicons corresponding to the cleavage events. The amplified segments from the cleavage reactions were also used as templates for NGS library preparation or as substrates for Sanger sequencing. Sequences with PAM activity compatible with the CRISPR complex were identified by sequencing the resulting library, which is a subset of the initial 8N library. For PAM testing using the treated RNA constructs, the same procedure was repeated except that in vitro transcription RNA was added along with the plasmid library, and the minimal CRISPR array template was omitted.
[0225] Analysis of the intergenetic regions surrounding the Cas effector and CRISPR array identified potential repeat repression sequences corresponding to the double-stranded tracrRNA sequence. TracrRNA and crRNA repeats were folded and trimmed, and the tetraloop sequence of GAAA was added to maintain the stem-loop region of the crRNA-tracrRNA complex.
[0226] Example 2A - In vitro targeted integrase activity Integrase activity was assayed using previously identified PAMs, but could instead be performed using PAM library substrates with reduced efficiency. One configuration of components for the in vitro test involved three plasmids other than the one containing the donor sequence: (1) an expression plasmid with an effector (or multiple effectors) under the T7 promoter, (2) an expression plasmid with a transposase gene under the T7 promoter, sgRNA or crRNA and tracrRNA, (3) a target plasmid containing a spacer site and appropriate PAMs, and (4) a donor plasmid containing the necessary left-end (LE) and right-end (RE) DNA sequences for transposition around the cargo gene (e.g., a selection marker such as a Tet resistance gene). Effector and transposase genes were expressed using an in vitro transcription / translation system (e.g., an E. coli lysate or reticulocyte lysate-based system). After expression, RNA, target DNA, and donor DNA were added and incubated to induce transposition. Transcription was detected via PCR across the transposase junction, with one primer on the target DNA and one on the donor DNA. The resulting PCR product was sequenced via NGS to determine the precise insertion topology to the sgRNA / crRNA target site. The primers were positioned downstream to accommodate and detect various insertion sites. The primers were designed so that insertion would be detected in either cargo orientation and on either side of the spacer, as the insertion direction had not been previously documented.
[0227] Similarly, integration efficiency was measured via qPCR of the experimental output of target DNA with integrated cargo, normalized to the amount of unmodified target DNA measured via quantitative PCR (qPCR).
[0228] This assay can be performed using purified protein components rather than from lysate-based expression. In this case, the protein was expressed in E. coli protease-deficient strain B under a T7-inducible promoter, the cells were lysed using sonication, and the His-tagged protein of interest was purified using Ni-NTA affinity chromatography on an FPLC system. Purity was determined using SDS-PAGE and concentration measurements of the degraded protein bands on a Coomassie-stained acrylamide gel. The protein was desalted in a storage buffer (or other buffer determined for maximum stability) consisting of pH 7.5, 50 mM Tris-HCl, 300 mM NaCl, 1 mM TCEP, and 5% glycerol, and stored at -80°C. After purification, the effector and transposase were added to the sgRNA, target DNA, and donor DNA described above in a reaction buffer, for example, 26 mM HEPES at pH 7.5 supplemented with 15 mM Mg(Oac)2, 4.2 mM TRIS at pH 8, 50 μg / mL BSA, 2 mM ATP, 2.1 mM DTT, 0.05 mM EDTA, 0.2 mM MgCl2, 28 mM NaCl, 21 mM KCl, and 1.35% glycerol (final pH 7.5).
[0229] Example 2B - In vitro activity Targeted nuclease In-situ expression and protein sequence analysis revealed that several RNA-induced effectors were active nucleases. They contained predicted endonuclease-related domains (matching the RuvC domain and HNH endonuclease domain), as well as / or predicted HNH and RuvC catalytic residues.
[0230] Candidate activity was tested using an E. coli lysate-based expression system and in vitro transcription RNA with engineered single guide RNA sequences. Active proteins that successfully cleaved the library produced a band of approximately 170 bp in the gel.
[0231] DNA integration and transposition A transposon is predicted to be active when the genomic sequence encoding the transposon contains one or more protein sequences with transposase and / or integrase function within the left and right ends of the transposon. The Tn7 transposon as defined herein may contain the catalytic transposase TnsB, but may also contain TnsA, TnsC, TnsD, TnsE, TniQ, and / or other transposases or integrases. The transposon terminus contains a predicted transposase binding site, including direct and / or inverted repeats of 15 bp–150 bp in length, adjacent to the transposase protein and other “cargo” genes. Protein sequence analysis shows that the transposase contains an integrase domain, a transposase domain, and / or transposase catalytic residues, suggesting that they are active (e.g., Figure 4A).
[0232] Target DNA integration Putative CRISPR-associated transposons (CASTs) include DNA and / or RNA-targeting CRISPR nucleases or effectors, as well as proteins with predicted transposase function located near CRISPR arrays. In some systems, the nuclease is predicted to be active based on the presence of an endonuclease-associated catalytic domain and / or catalytic residues.
[0233] In some systems, the effector is predicted to have homology to the documented CRISPR effector protein but be inactive based on the absence of the endonuclease domain and / or catalytic residues. The transposase is predicted to be associated with the effector when the CRISPR locus (inactive CRISPR nuclease and array) and the transposase protein are located within the left and right ends of the predicted transposon (Figure 4A). In this case, the effector is predicted to induce DNA integration at a specific genomic location based on guide RNA.
[0234] CAST activity was tested using five types of components: (1) Cas effector protein expressed by an in vitro expression system, (2) target DNA fragment or plasmid containing the target sequence and PAM corresponding to the Cas enzyme, (3) donor DNA fragment containing DNA markers or fragments adjacent to the LE and RE of the transposase system in the DNA fragment or plasmid, (4) any combination of transposase proteins expressed using an in vitro expression system, and (5) an engineered in vitro transcriptional single guide RNA sequence. Active systems with successful donor fragment transposition were assayed by PCR amplification of the donor-target junction.
[0235] After the rearrangement reaction, PCR amplification of the junction showed that appropriate donor-target formation occurred, indicating that the rearrangement reaction was sg-dependent (Figure 6). PCR amplification of reactions #3 and #4 showed that both donor orientations toward the target occurred, i.e., LE was more closely aligned with PAM, and RE was more closely aligned with PAM. Although both rearrangement orientations occurred, there was a preference for donor incorporation at the target where LE was closer to PAM, as represented by the strong bands present for reactions #4 and #5.
[0236] Sanger sequencing was performed on the preferred orientation products. Among the integrations where the LE was closer to the PAM, there was a clear degradation of the sequencing chromatogram signal from either the forward or reverse direction across the target / donor junction. This indicated that among the products where the LE was oriented closer to the PAM, the integration occurred at a certain range of nucleotides, with the major product among the PAM-closer LE products being a 61 bp integration from the PAM (Figure 7A). Sequencing from the donor across the donor-target junction defined the composition of the essential outer boundary lines of the LE and RE sequences (Figures 7A-7B). Further investigation of the LE and RE domains may determine the inner limits of the LE and RE sequences essential for the transposition. Sequencing of the RE on the PAM-closer LE product showed a 3 bp duplication downstream of the donor RE (Figure 7B). This is partly due to the Tn7 transposase integration event, which cleaved and ligated the donor fragment at alternating cleavage sites. The 3bp duplication is smaller than the expected 5bp duplication from other Tn7 transposases.
[0237] Sanger sequencing of PCR amplification products of the target plasmid against an 8N library also revealed the PAM preference of the MG64-1 effector as nGTn / nGTt on the 5' end of the spacer (Figure 7C). NGS analysis of the PAM library target confirmed the nGTn motif specificity at the 5' end.
[0238] Example 3 - Predicted RNA folding Using the Andronescu 2007 method, the predicted RNA folding of the active single RNA sequence was calculated at 37°. All hairpin loop secondary structures were individually removed from the structure and compiled by iterating into a smaller single guide. In a second approach, the tracrRNA of MG64-1 was aligned to a documented Vk-type tracrRNA, and the region of the unique insertion was mutated from the single guide and minimized by 57 bases. Figure 12A depicts the predicted structure of MG64-1 sgRNA. Figure 12B depicts the predicted structure of MG64-3 sgRNA. Figure 12C depicts the predicted structure of MG64-5 sgRNA. The color of the bases corresponds to the probability of base pairing for that base, with red representing a high probability and blue representing a low probability.
[0239] Example 4 - Transposon end verification via gel shift Transposon ends were tested for TnsB binding via electrophoretic mobility shift assay (EMSA). In this case, potential LE or RE were synthesized as DNA fragments (100-500 bp) and end-labeled with FAM via PCR using FAM-labeled primers. TnsB proteins were synthesized in an in vitro transcription / translation system. After synthesis, 1 μL of TnsB protein was added to 50 nM labeled RE or LE in 10 μL of reaction in binding buffer (20 mM HEPES at pH 7.5, 2.5 mM Tris at pH 7.5, 10 mM NaCl, 0.0625 mM EDTA, 5 mM TCEP, 0.005% BSA, 1 ug / mL poly(dI-dC), and 5% glycerol). The binding reaction was incubated at 30°C for 40 minutes, followed by the addition of 2 μL of 6X loading buffer (60 mM KCl, 10 mM Tris at pH 7.6, 50% glycerol). The binding reaction product was separated on a 5% TBE gel and visualized. A shift in LE or RE in the presence of TnsB indicated successful binding and transposase activity (Figure 24).
[0240] Example 5 - Integrase activity in E. coli Because E. coli lacks the ability to efficiently repair double-strand DNA breaks in its genome, transformation of E. coli with drugs capable of inducing double-strand breaks in the E. coli genome leads to cell death. This phenomenon was utilized to test endonuclease or effector-assisted integrase activity in E. coli by recombinantly expressing either an endonuclease or an effector-assisted integrase, along with a guide RNA (determined, e.g., as in Example 3), in target strains having integrated spacer / target and PAM sequences in their genomic DNA.
[0241] Next, the manipulated strains were transformed with plasmids containing a nuclease or effector with a single guide RNA, plasmids expressing integrase and accessory genes, and plasmids containing a temperature-sensitive origin of replication with selectable markers adjacent to left- and right-hand transposon motifs for integration. The transformants induced for the expression of these genes were then screened for marker migration to genomic targets by selection at limiting temperatures for plasmid replication, and marker integration within the genome was confirmed by PCR.
[0242] An unbiased approach was used to screen for off-target insertions. Briefly, purified gDNA was fragmented by Tn5 transposase or shear, and the DNA of interest was then PCR-amplified using primers specific to ligated adapters and selectable markers. The amplicons were then prepared for NGS sequencing. Analysis of the resulting sequences was performed, trimming from the transposon sequences, and adjacent sequences were mapped to the genome to determine insertion sites and the off-target insertion rate.
[0243] Example 6 - Colony PCR screening of transposase activity For testing nuclease or effector-assisted integrase activity in bacterial cells, strain MGB0032 was constructed from BL21(DE3)E. coli cells engineered to contain MG64_1-specific targets and corresponding PAM sequences. MGB0032 E. coli cells were then transformed with pJL56 (a plasmid expressing the MG64_1 effector and helper suite, ampicillin-resistant) and pTCM 64_1 sg, a chloramphenicol-resistant plasmid expressing a single guide RNA sequence of the engineered target of interest, driven by the T7 promoter.
[0244] Next, MGB0032 cultures containing both plasmids were grown to saturation and diluted at least 1:10 in growth cultures containing appropriate antibiotics, and incubated at 37°C to approximately 1 OD. Cells from this growth stage were electrocompetent and transformed with streamlined 64_1 pDonor, a plasmid with tetracycline resistance markers flanked by left-end (LE) and right-end (RE) transposon motifs for integration. The electroporated cells were then plated onto LB-agar-ampicillin-chloramphenicol-tetracycline and harvested on LB medium for 2 hours in or without IPTG at a final concentration of 100 μM before incubation at 37°C for 4 days. Using a sterile toothpick, the resulting CFUs were sampled and mixed in water. To this solution, Q5 High Fidelity PCR mastermix (New England Biolabs) and primers LA155 (5'-GCTCTTCCGATCTNNNNNGATGAGCGCATTGTTAGATTTCAT-3' (SEQ ID NO: 1256)) and oJL50 (5'-AAACCGACATCGCAGGCTTC-3' (SEQ ID NO: 1257)) were added. These primers are adjacent to the predicted insertion junction. The predicted product size was 609 bp. The DNA amplified PCR product was visualized on a 2% agarose gel. Sanger sequencing of the PCR product confirmed the transposition event.
[0245] Example 7 - Intracellular Expression / In Vitro Assay To test the functionality of NLS constructs in physiologically relevant environments, constructs cloned with active NLS-tagged CAST components were incorporated into K562 cells using lentiviral transduction. Briefly, constructs cloned into lentiviral transfer plasmids were transfected into 293T cells using envelope and packaging plasmids, and the virus, including the supernatant, was collected from the medium after 72 hours of incubation. The virus-containing medium was then incubated with K562 cell lines containing 8 μg / mL of polyblen for 72 hours, and then selected for collective incorporation of transfected cells using 1 μg / mL of puromycin for 4 days. The selected cell lines were collected at the end of the 4 days and differentially lysed for nuclear and cytoplasm...
Claims
1. A system for rearranging a cargo nucleotide sequence into a target nucleic acid site within a target nucleic acid, a) A Cas effector complex comprising a class 2 V-type Cas effector, a microprokaryotic ribosomal protein subunit S15, and an engineered guide polynucleotide that hybridizes to the target nucleic acid site, b) A Tn7 type transposase complex comprising an auxiliary protein that binds to the Cas effector complex and includes a sequence having at least 70% sequence homology with the TnsB, TnsC, and TniQ components, as well as one of SEQ ID NOs. 228-230 and 235-249, c) A system comprising a double-stranded nucleic acid that interacts with the Tn7 type transposase complex and contains the cargo nucleotide sequence.
2. The system according to claim 1, wherein the Cas effector complex is non-covalently bonded to the Tn7 type transposase complex.
3. The system according to claim 1, wherein the Cas effector complex is covalently bonded to the Tn7 type transposase complex.
4. The system according to claim 1, wherein the Cas effector complex is fused to the Tn7 type transposase complex.
5. The system according to any one of claims 1 to 4, wherein the cargo nucleotide sequence is adjacent to a left transposase recognition sequence and a right transposase recognition sequence recognized by the Tn7 type transposase complex.
6. The system according to claim 5, wherein the left transposase recognition sequence includes a sequence having at least 80% identity with any one of sequence numbers 9, 11, 36-38, 76, and 78.
7. The system according to claim 5, wherein the right transposase recognition sequence includes a sequence having at least 80% identity with any one of sequence numbers 8, 10, 39-44, 77, 79, and 93.
8. The system according to any one of claims 1 to 7, wherein the target nucleic acid includes a PAM sequence compatible with the Cas effector complex.
9. The system according to claim 8, wherein the PAM sequence includes sequence number 31.
10. The system according to claim 8 or 9, wherein the PAM sequence is located about 50 to about 70 base pairs from the target nucleic acid site.
11. The system according to claim 10, wherein the PAM sequence is located at 3' of the target nucleic acid site.
12. The system according to claim 10, wherein the PAM sequence is located at 5' of the target nucleic acid site.
13. The system according to any one of claims 1 to 12, wherein the V-type Cas effector of class 2 is a Cas 12k effector.
14. The system according to any one of claims 1 to 12, wherein the Class 2 V-type Cas effector comprises a polypeptide having at least 80% identity with any one of sequence numbers 1, 12, 16, 20-30, 64, 80-85, and 220.
15. The system according to any one of claims 1 to 12, wherein the Class 2 V-type Cas effector comprises a polypeptide having a sequence having at least 90% identity with any one of sequence numbers 1, 12, 16, 20-30, 64, 80-85, and 220.
16. The system according to any one of claims 1 to 12, wherein the Class 2 V-type Cas effector comprises a polypeptide having any one sequence from sequence numbers 1, 12, 16, 20-30, 64, 80-85, and 220.
17. The system according to any one of claims 1 to 16, wherein the TnsB component comprises a polypeptide having a sequence that is at least 80% identical to any one of sequence numbers 2, 13, 17, and 65.
18. The system according to any one of claims 1 to 16, wherein the TnsB component comprises a polypeptide having a sequence that is at least 90% identical to any one of sequence numbers 2, 13, 17, and 65.
19. The system according to any one of claims 1 to 16, wherein the TnsB component comprises a polypeptide having any one sequence of sequence numbers 2, 13, 17, and 65.
20. The system according to any one of claims 1 to 19, wherein the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 80% identity with any one of sequence numbers 3-4, 14-15, 18-19, 66-67, and 109-111.
21. The system according to any one of claims 1 to 19, wherein the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently containing a sequence having at least 90% identity with any one of sequence numbers 3-4, 14-15, 18-19, 66-67, and 109-111.
22. The system according to any one of claims 1 to 19, wherein the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising one sequence from among sequence numbers 3-4, 14-15, 18-19, 66-67, and 109-111.
23. The system according to any one of claims 1 to 22, wherein the manipulated guide polynucleotide comprises a sequence containing at least 46 to 80 consecutive nucleotides having at least 80% identity with any one of SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222.
24. The system according to any one of claims 1 to 22, wherein the manipulated guide polynucleotide includes a sequence having at least 80% sequence homology with any one of sequence numbers 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944.
25. The system according to any one of claims 1 to 23, wherein the microprokaryotic ribosomal protein subunit S15 includes a sequence having at least 80% sequence homology with any one of sequence numbers 187 to 189.
26. The system according to any one of claims 1 to 23, wherein the microprokaryotic ribosomal protein subunit S15 is encoded by a sequence having at least 80% sequence homology with any one of sequence numbers 181 to 183.
27. The system according to any one of claims 1 to 26, wherein the class 2 V-type Cas effector and the Tn7-type transposase complex are encoded by a polynucleotide sequence containing less than about 10 kilobases.
28. The system according to any one of claims 1 to 27, wherein the auxiliary protein comprises a sequence having at least 80% sequence homology with any one of sequence numbers 235 to 249.
29. A system for rearranging a cargo nucleotide sequence into a target nucleic acid site within a target nucleic acid, a) A Cas effector complex comprising a class 2 V-type Cas effector, a microprokaryotic ribosomal protein subunit S15, and an engineered guide polynucleotide that hybridizes to the target nucleic acid site, b) A Tn7 type transpotase complex that binds to the Cas effector complex and includes a functional domain (FD)-TniQ fusion and an auxiliary protein, c) A system comprising a double-stranded nucleic acid that interacts with the Tn7 type transposase complex and contains the cargo nucleotide sequence.
30. The system according to claim 29, wherein the Cas effector complex is non-covalently bonded to the Tn7 type transposase complex.
31. The system according to claim 29, wherein the Cas effector complex is covalently bonded to the Tn7 type transposase complex.
32. The system according to claim 29, wherein the Cas effector complex is fused to the Tn7 type transposase complex.
33. The system according to claim 29, wherein the functional domain (FD) includes a sequence having at least 80% sequence homology with any one of sequence numbers 257-307 and 1138-1242.
34. The system according to any one of claims 29 to 33, wherein the cargo nucleotide sequence is adjacent to a left transposase recognition sequence and a right transposase recognition sequence recognized by the Tn7 type transposase complex.
35. The system according to claim 34, wherein the left-side transposase recognition sequence includes a sequence having at least 80% identity with any one of sequence numbers 9, 11, 36-38, 76, and 78.
36. The system according to any one of claims 34, wherein the right-side transposase recognition sequence includes a sequence having at least 80% identity with any one of sequence numbers 8, 10, 39-44, 77, 79, and 93.
37. The system according to any one of claims 29 to 36, wherein the target nucleic acid includes a PAM sequence compatible with the Cas effector complex.
38. The system according to claim 37, wherein the PAM sequence includes sequence number 31.
39. The system according to claim 37 or 38, wherein the PAM sequence is located about 50 to about 70 base pairs from the target nucleic acid site.
40. The system according to claim 39, wherein the PAM sequence is located at 3' of the target nucleic acid site.
41. The system according to claim 40, wherein the PAM sequence is located at 5' of the target nucleic acid site.
42. The system according to any one of claims 29 to 41, wherein the V-type Cas effector of class 2 is a Cas 12k effector.
43. The system according to any one of claims 29 to 42, wherein the Class 2 V-type Cas effector comprises a polypeptide having a sequence having at least 90% identity with any one of sequence numbers 1, 12, 16, 20-30, 64, 80-85, and 220.
44. The system according to any one of claims 29 to 43, wherein the Class 2 V-type Cas effector comprises a polypeptide having any one sequence from sequence numbers 1, 12, 16, 20-30, 64, 80-85, and 220.
45. The system according to any one of claims 29 to 44, wherein the manipulated guide polynucleotide comprises a sequence containing at least 46 to 80 consecutive nucleotides having at least 80% identity with any one of SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222.
46. The system according to any one of claims 29 to 45, wherein the manipulated guide polynucleotide comprises a sequence having at least 80% sequence homology with any one of sequence numbers 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944.
47. The system according to any one of claims 29 to 46, wherein the microprokaryotic ribosomal protein subunit S15 includes a sequence having at least 80% sequence homology with any one of sequence numbers 187 to 189.
48. The system according to any one of claims 29 to 46, wherein the microprokaryotic ribosomal protein subunit S15 is encoded by a sequence having at least 80% sequence homology with any one of sequence numbers 181 to 183.
49. The system according to any one of claims 29 to 48, wherein the class 2 V-type Cas effector and the Tn7-type transposase complex are encoded by a polynucleotide sequence of less than about 10 kilobases.
50. The system according to any one of claims 29 to 49, wherein the auxiliary protein comprises a sequence having at least 70% sequence homology with any one of sequence numbers 228 to 230 and 235 to 249.
51. The system according to claim 50, wherein the auxiliary protein is ClpX containing a sequence having at least 80% sequence homology with any one of sequence numbers 235 to 249.
52. A system for rearranging a cargo nucleotide sequence into a target nucleic acid site within a target nucleic acid, a) A Cas effector complex comprising a class 2 V-type Cas effector and an engineered guide polynucleotide that hybridizes to the target nucleic acid site, wherein the Cas effector complex comprises a polypeptide having at least 80% sequence homology with any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220, b) A Tn7 type transposase complex comprising a Tn7 type transposase complex that binds to the Cas effector complex and comprises TnsB, TnsC, and TniQ components, wherein the TnsB, TnsC, or TniQ component contains a sequence having at least 80% sequence homology with any one of SEQ ID NOs: 2-4, 13-15, 17-19, 65-67, and 109-111; and a Tn7 type transposase complex comprising an auxiliary protein containing a sequence having at least 70% sequence homology with any one of SEQ ID NOs: 228-230 and 235-249, c) A system comprising a double-stranded nucleic acid that interacts with the Tn7 type transposase complex and contains the cargo nucleotide sequence.
53. The system according to claim 29, wherein the Cas effector complex is non-covalently bonded to the Tn7 type transposase complex.
54. The system according to claim 29, wherein the Cas effector complex is covalently bonded to the Tn7 type transposase complex.
55. The system according to claim 29, wherein the Cas effector complex is fused to the Tn7 type transposase complex.
56. The system according to any one of claims 29 to 55, wherein the cargo nucleotide sequence is adjacent to a left transposase recognition sequence and a right transposase recognition sequence recognized by the Tn7 type transposase complex.
57. The system according to claim 56, wherein the left-side transposase recognition sequence includes a sequence having at least 80% identity with any one of sequence numbers 9, 11, 36-38, 76, and 78.
58. The system according to claim 56, wherein the right-side transposase recognition sequence includes a sequence having at least 80% identity with any one of sequence numbers 8, 10, 39-44, 77, 79, and 93.
59. The system according to any one of claims 29 to 58, wherein the target nucleic acid includes a PAM sequence compatible with the Cas effector complex.
60. The system according to claim 59, wherein the PAM sequence includes sequence number 31.
61. The system according to claim 59 or 60, wherein the PAM sequence is located about 50 to about 70 base pairs from the target nucleic acid site.
62. The system according to claim 61, wherein the PAM sequence is located at 3' of the target nucleic acid site.
63. The system according to claim 61, wherein the PAM sequence is located at 5' of the target nucleic acid site.
64. The system according to any one of claims 29 to 63, wherein the V-type Cas effector of class 2 is a Cas 12k effector.
65. The system according to any one of claims 29 to 63, wherein the Class 2 V-type Cas effector comprises a polypeptide having a sequence having at least 90% identity with any one of sequence numbers 1, 12, 16, 20-30, 64, 80-85, and 220.
66. The system according to any one of claims 29 to 63, wherein the Class 2 V-type Cas effector comprises a polypeptide having one of the sequences 1, 12, 16, 20-30, 64, 80-85, and 220.
67. The system according to any one of claims 29 to 63, wherein the TnsB, TnsC, or TniQ component includes a sequence having at least 90% sequence homology with any one of sequence numbers 2 to 4, 13 to 15, 17 to 19, 65 to 67, and 109 to 111.
68. The system according to any one of claims 29 to 63, wherein the TnsB, TnsC, or TniQ component includes one of the sequences 2 to 4, 13 to 15, 17 to 19, 65 to 67, and 109 to 111.
69. The system according to any one of claims 29-68, wherein the manipulated guide polynucleotide comprises a sequence containing at least 46 to 80 consecutive nucleotides having at least 80% identity with any one of SEQ ID NOs. 5-6, 32-33, 94-95, 104-105, 119-122, and 222.
70. The system according to any one of claims 29 to 68, wherein the manipulated guide polynucleotide comprises a sequence having at least 80% sequence homology with any one of sequence numbers 106, 107, 108, 5, 45-63, 68-75, 96-103, 123-140, and 754-944.
71. The system according to any one of claims 29 to 70, wherein the microprokaryotic ribosomal protein subunit S15 includes a sequence having at least 80% sequence homology with any one of sequence numbers 187 to 189.
72. The system according to any one of claims 29 to 70, wherein the microprokaryotic ribosomal protein subunit S15 is encoded by a sequence having at least 80% sequence homology with any one of sequence numbers 181 to 183.
73. The system according to any one of claims 29 to 72, wherein the class 2 V-type Cas effector and the Tn7-type transposase complex are encoded by a polynucleotide sequence of less than about 10 kilobases.
74. The system according to any one of claims 29 to 73, wherein the auxiliary protein is ClpX containing a sequence having at least 80% sequence homology with any one of sequence numbers 235 to 249.
75. A system for rearranging a cargo nucleotide sequence into a target nucleic acid site within a target nucleic acid, a) A Cas effector complex that hybridizes to the target nucleic acid site, i) A Class 2 V-type Cas effector comprising a sequence having at least 80% sequence homology with any one of sequence numbers 1, 81, 82, 83, and 85, and ii) A Cas effector complex comprising an engineered guide polynucleotide having at least 80% identity with any one of sequence numbers 5, 6, 45-63, 68-75, 96-103, 123-140, and 754-944, b) A Tn7 type transposase complex that binds to the Cas effector complex and comprises TnsB, TnsC, and TniQ components, and a Tn7 type transposase complex wherein the TnsB, TnsC, or TniQ component contains a sequence having at least 80% sequence homology with any one of SEQ ID NOs. 2 to 4, and the auxiliary protein contains a sequence having at least 70% sequence homology with any one of SEQ ID NOs. 228 to 230 and 235 to 249, c) A double-stranded nucleic acid that interacts with the Tn7 type transposase complex and is in the order of 5' to 3', i) A left transposase recognition sequence comprising a sequence having at least 80% sequence homology with any one of sequence numbers 9, 11, 36, 37, and 38, ii) The cargo nucleotide sequence and, ii) A system comprising a double-stranded nucleic acid, which comprises a right-side transposase sequence having at least 80% identity with any one of sequence numbers 8, 39-44, and 93.
76. A system for rearranging a cargo nucleotide sequence into a target nucleic acid site within a target nucleic acid, a) A Cas effector complex that hybridizes to the target nucleic acid site, i) A Class 2 V-type Cas effector comprising a sequence having at least 80% sequence homology with sequence number 12, and iii) A Cas effector complex comprising an engineered guide polynucleotide having at least 80% identity with any one of sequence numbers 32, 102, 104, and 107, b) A Tn7 type transposase complex that binds to the Cas effector complex and includes TnsB, TnsC, and TniQ components, wherein the TnsB, TnsC, or TniQ component includes a sequence having at least 80% sequence homology with any one of SEQ ID NOs. 13 to 15, and the auxiliary protein has a sequence having at least 70% sequence homology with any one of SEQ ID NOs. 228 to 230 and 235 to 249, c) A double-stranded nucleic acid that interacts with the Tn7 type transposase complex and is in the order of 5' to 3', i) A left-side transposase sequence containing a sequence having at least 80% sequence homology with sequence number 76, ii) The cargo nucleotide sequence, and iii) A system comprising a double-stranded nucleic acid, which includes a right-side transposase sequence having at least 80% identity with sequence number 77.
77. A system for rearranging a cargo nucleotide sequence into a target nucleic acid site within a target nucleic acid, a) A Cas effector complex that hybridizes to the target nucleic acid site, i) A Class 2 V-type Cas effector comprising a sequence having at least 80% sequence homology with sequence number 16, and ii) A Cas effector complex comprising an engineered guide polynucleotide having at least 80% identity with any one of sequence numbers 33, 103, 105, and 108, b) A Tn7 type transposase complex that binds to the Cas effector complex and includes TnsB, TnsC, and TniQ components, wherein the TnsB, TnsC, or TniQ component includes a sequence having at least 80% sequence homology with any one of SEQ ID NOs. 17-19, and the auxiliary protein includes a sequence having at least 70% sequence homology with any one of SEQ ID NOs. 228-230 and 235-249, c) A double-stranded nucleic acid that interacts with the Tn7 type transposase complex and is in the order of 5' to 3', i) A left transposase recognition sequence containing a sequence having at least 80% sequence homology with sequence number 78, ii) The cargo nucleotide sequence, and iii) A system comprising a double-stranded nucleic acid, which includes a right-side transposase sequence having at least 80% identity with sequence number 79.
78. The system according to any one of claims 75 to 77, further comprising a PAM sequence compatible with the Cas effector complex.
79. The system according to claim 78, wherein the PAM sequence includes sequence number 31.
80. The system according to claim 78 or 79, wherein the PAM sequence is located about 50 to about 70 base pairs from the target nucleic acid site.
81. The system according to claim 80, wherein the PAM sequence is located at 3' of the target nucleic acid site.
82. The system according to claim 80, wherein the PAM sequence is located at 5' of the target nucleic acid site.
83. The system according to any one of claims 75 to 82, wherein the Cas effector complex further comprises the microprokaryotic cell ribosomal protein subunit S15.
84. The system according to claim 83, wherein the microprokaryotic ribosomal protein subunit S15 includes a sequence having at least 80% sequence homology with any one of sequence numbers 187 to 189.
85. The system according to any one of claims 75 to 84, wherein the auxiliary protein is ClpX containing a sequence having at least 80% sequence homology with any one of sequence numbers 235 to 249.
86. A manipulated nuclease system, a) An endonuclease containing a RuvC domain, derived from an uncultured microorganism, and a class 2 V-K type Cas effector having at least 80% identity with any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220, b) An engineered nuclease system comprising an engineered guide RNA, the engineered guide polynucleotide comprising a sequence having at least approximately 80% identity with one of sequence numbers 754-944, wherein the engineered guide polynucleotide comprises a sequence that forms a complex with an endonuclease and hybridizes to a target nucleic acid sequence.
87. A method for transposing a cargo nucleotide sequence into a target nucleic acid site, comprising introducing the system described in any one of claims 1 to 86 into a cell.
88. A cell comprising the system described in any one of claims 1 to 86.
89. The cell according to claim 88, wherein the cell is a eukaryotic cell.
90. The cell according to claim 88, wherein the cell is a mammalian cell.
91. The cell according to claim 88, wherein the cell is an immortalized cell.
92. The cell according to claim 88, wherein the cell is an insect cell.
93. The cell according to claim 88, wherein the cell is a yeast cell.
94. The cell according to claim 88, wherein the cell is a plant cell.
95. The cell according to claim 88, wherein the cell is a fungal cell.
96. The cell according to claim 88, wherein the cell is a prokaryotic cell.
97. The cell according to claim 88, wherein the cell is A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cell, HT1080, HepG2, Huh7, K562, primary cell, or a derivative thereof.
98. The cell according to claim 88, wherein the cell is an engineered cell.
99. The cell according to claim 88, wherein the cell is a stable cell.