Systems and methods for transposing cargo nucleotide sequences

CRISPR-associated transposases (CAST) enhance gene editing by improving integration efficiency and specificity, addressing the limitations of existing technologies to treat genetic diseases through targeted DNA cargo delivery in diverse cell types.

WO2026035770A1PCT designated stage Publication Date: 2026-02-12METAGENOMI INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/040776
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-05-09
Filing Date
2025-08-05
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Current gene editing technologies face challenges in efficiently and specifically integrating large transgenes into the human genome, particularly in cell types where DNA repair machinery may not be active, limiting their therapeutic potential for diverse genetic diseases.

Method used

Repurposing CRISPR-associated transposases (CAST) from microbes for site-specific, targeted integration of gene-sized DNA cargoes, enhanced by protein engineering and optimized delivery systems, including an 'all-in-one' mRNA approach to facilitate efficient integration in various cell types, especially primary hepatocytes, without relying on DNA breaks.

Benefits of technology

The CAST system achieves >50x improvement in integration efficiency compared to natural systems, enabling treatment of genetic diseases caused by loss-of-function mutations and advancing biotechnology and synthetic biology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025040776_12022026_PF_FP_ABST
    Figure US2025040776_12022026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides systems and methods for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid. These systems and methods may comprise a double-stranded nucleic acid comprising the cargo nucleotide sequence, wherein the cargo nucleotide sequence interacts with a transposase recognition complex, an effector complex comprising an effector and at least one engineered guide polynucleotide that hybridizes to the target nucleic acid, and the transposase recognition complex wherein the transposase recognition complex recruits the cargo nucleotide to the target nucleic acid site.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. MTGNZ000100WOSYSTEMS AND METHODS FOR TRANSPOSING CARGO NUCLEOTIDE SEQUENCESCROSS-REFERENCE

[0001] This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 679,302 filed August 5, 2024, and U.S. Provisional Patent Application No. 63 / 803,019 filed May 9, 2025, each of which is incorporated by reference in its entirety herein.SEQUENCE LISTING

[0002] The instant application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on August 5, 2025, is named 04_MTGNZ00100WC)_20250805_sequence_listing.xml and is 1,623,797 bytes in size.BACKGROUND

[0003] Gene editing technologies capable of efficient and specific integration of large transgenes into the human genome have the potential to bring curative treatments to diverse patient populations by providing a single approach that could cover a majority of diseasecausing mutations.SUMMARY

[0004] CRISPR-associated transposases (CAST) discovered from microbes may be repurposed and programmed for site-specific, targeted integration of gene-sized DNA cargoes into the human genome. CAST protein engineering and delivery optimization enable efficient integration of diverse cargo designs across different cell types and target loci, including with therapeutically -relevant delivery approaches that are expected to be required for in vivo editing. Through structure and protein design, integration efficiency of a compact type V-K CAST system derived from uncultivated microbes may be improved by >50x compared with the natural system. Furthermore, an ‘all-in-one’ mRNA expressing multiple protein components may be used to deliver the system for efficient editing in primary hepatocytes. Combined with the fact that CAST integration uses a transposase to circumvent doublestranded DNA breaks, these advancements suggest important advantages over other gene editing approaches, especially those that are dependent on DNA repair machinery that may not be active in cell types targeted in vivo. Given their programmability and efficiency integrating large DNA cargoes, CAST systems may have the potential to treat any genetic disease caused by a loss-of-function mutation, and contribute to other important advances in biotechnology and synthetic biology.Attorney Docket No. MTGNZOOOIOOWO

[0005] Described herein, in certain embodiments, are systems for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid comprising: a) a Cas effector complex and a Tn7 type transposase complex that binds the Cas effector complex, wherein the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence having at least 70% sequence identity to any one of SEQ ID NOs: 1406- 1409; and b) a double-stranded nucleic acid that interacts with the Tn7 type transposase complex and comprises the cargo nucleotide sequence. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence of any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex binds non-covalently to the Tn7 type transposase complex. In some embodiments, the Cas effector complex is covalently linked to the Tn7 type transposase complex. In some embodiments, the Cas effector complex is fused to the Tn7 type transposase complex. In some embodiments, the cargo nucleotide sequence is flanked by a left-hand transposase recognition sequence and a right-hand transposase recognition sequence recognized by the Tn7 type transposase complex. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least 80% identity to SEQ ID NO: 1404. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least 80% identity to SEQ ID NO: 1405. In some embodiments, the engineered guide polynucleotide comprises a nucleotide sequence comprising at least about 46-80 consecutive nucleotides having at least 80% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide comprises a nucleotide sequence comprising having at least 80% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411.

[0006] Described herein, in certain embodiments, are systems for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid comprising: a) a Cas effector complex and a Tn7 type transposase complex that binds the Cas effector complex, wherein the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence having at least 70% sequence identity to any one of SEQ ID NOs: 1406- 1409; and b) a double- stranded nucleic acid that interacts with the Tn7 type transposase complex and comprising in 5’ to 3’ order: i) a left-hand transposase recognition sequence comprising a nucleotide sequence having at least 80% sequence identity to SEQ ID NO: 1404;Attorney Docket No. MTGNZ000100WO ii) the cargo nucleotide sequence; and iii) a right-hand transposase recognition sequence comprising a nucleotide sequence having at least 80% identity to SEQ ID NO: 1405. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1406- 1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence of any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex binds non-covalently to the Tn7 type transposase complex. In some embodiments, the Cas effector complex is covalently linked to the Tn7 type transposase complex. In some embodiments, the Cas effector complex is fused to the Tn7 type transposase complex. In some embodiments, the cargo nucleotide sequence is flanked by a left-hand transposase recognition sequence and a right-hand transposase recognition sequence recognized by the Tn7 type transposase complex. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least 80% identity to SEQ ID NO: 1404. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least 80% identity to SEQ ID NO: 1405. In some embodiments, the engineered guide polynucleotide comprises a nucleotide sequence comprising at least about 46-80 consecutive nucleotides having at least 80% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide comprises a nucleotide sequence comprising having at least 80% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411.

[0007] Described herein, in certain embodiments, are methods for transposing a cargo nucleotide sequence into a target nucleic acid site comprising introducing the systems described herein to a cell.

[0008] Described herein, in certain embodiments, are cells comprising the systems described herein. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is an immortalized cell. In some embodiments, the cell is an insect cell. In some embodiments, the cell is a yeast cell. In some embodiments, the cell is a plant cell. In some embodiments, the cell is a fungal cell. In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is an A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Veto, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cell, HT1080, HepG2, Huh7, K562, primary cell, or a derivative thereof. In some embodiments, the cell is an engineered cell. In some embodiments, the cell is a stable cell.Attorney Docket No. MTGNZ000100WO

[0009] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The novel features of the disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:

[0011] FIGs. 1A and IB depict on- and off- targeting integration profiles for MG64-1 and MG64-6 at four E. coli loci. FIG. 1A shows charts illustrating the number of integration events represented by single vs. co-integration events. FIG. 1A shows a bar plot illustrating percentage of reads mapping to an integration event targeting the engineered or endogenous targets in the forward (T-LR) vs. reverse (T-RL) orientation. No events were detected for the T-RL orientation in these experiments

[0012] FIGs. 2A and 2B depict MG64-1 Casl2k effector fusions for activity in human cells. FIG. 2A shows a schematic representation for in vitro testing of nuclear extracts to determine functionality of CAST components post-localization. Cells expressing components of interest are differentially lysed to yield cytoplasmic and nuclear extracts. Upon incubation with the addition of sgRNA, donor fragment and pTarget, PCR junctions are amplified. FIG. 2B shows images of PCR results illustrating integration junction PCR products of in vitro transposition reactions using nuclear extract inputs from cells expressing Casl2k-NLS with (lanes 3-6) and without (lanes 8-10) sso7d, NLS-TnsB, NLS-TnsC and either NLS-HMGNl-TniQ or NLS- Hlcore-TniQ. In vitro control reactions (lanes 1-2) contain all components expressed exclusively in vitro and tested in vitro. Expected bands with added sgRNA (+) indicate positive integration (arrows).

[0013] FIGs. 3A and 3B illustrate that nuclear localization tags and fusion domains for mammalian cell cargo integration are system-specific depicts the two pathways found in Tn7 and Tn7-like elements. FIG. 3A depicts a schematic of plasmids described herein. The “All-in- one pHelper no sg” plasmid is used as a negative control. FIG. 3B shows a bar plot illustratingAttorney Docket No. MTGNZ000100WO on-target integration efficiency for the expected LE integration junction product for MG64-1 and ShCAST determined from NGS sequencing. Quantified integration efficiencies are shown as the mean (bar) of four biological replicates (dots) with one standard deviation (whisker). Null represents the non-targeting spacer (negative control, dark grey dots), while T1 represents the AAVS1 target 1 (light grey dots).

[0014] FIGs. 4A and 4B illustrate integration efficiency measured by amplicon sequencing (NGS) of the left end target-to-donor junction for MG64-1 variants at AAVS1 in HEK293T cells (FIG. 4A) or Hep3B cells (FIG. 4B). Integration was measured after delivering via lipofection a plasmid containing all CAST protein components and a donor plasmid containing a full factor VIII gene (~4.7 kb in length) and the targeting sgRNA.

[0015] FIG. 5 illustrates integration efficiency measured by amplicon sequencing (NGS) of the left end target-to-donor junction for MG64-1 variants at AAVS1 in HEK293T cells, without ClpX. Integration was measured after delivering via lipofection a plasmid containing all CAST protein variants encoded in different order (C = TnsC; B = TnsB; Q = TniQ; KS = Casl2k-ASR-2A-S15; KS-WT = WT Casl2k-2A-S15; B-WT = WT TnsB). The donor plasmid containing a full factor VIII gene (~4.7 kb in length) and the targeting sgRNA were delivered in replicating or non-replicating plasmids.

[0016] FIG. 6 illustrates an all-in-one (AIO) mRNA encoding improved CAST protein variants was delivered, with or without ClpX mRNA and donor DNA plasmid containing factor VIII (gene of interest, GOI), simultaneously or in a 6-hour interval lipofection reaction (staggered), at a 3: 1 or 6.8:1.8 dose ratio. Staggered transfection: mRNA delivered 30 minutes after pDonor delivery. Integration efficiency in PMH cells was measured by NGS amplification of the target-to-left end junction product.

[0017] FIG. 7 illustrates a schematic of plasmids used in mammalian cell experiments. FIG. 7A is a schematic of plasmids used for screening integration activity in human immortalized HEK293T cells. FIG. 7B is a schematic of the all-in-one (AIO) mRNA and donor plasmid used for screening integration in human immortalized Hepal-6 cells.

[0018] FIG. 8 illustrates how the TnsB engineered variant R347A improves MG64-1 integration by >7 fold. Integration efficiency in HEK293T cells with MG64-1 TnsB engineered variants based on deep mutational scanning of 15 selected sites. Variants were screened in HEK293T cells with (FIG. 8A) and without (FIG. 8B) addition of a plasmid containing ClpX. Integration efficiency was measured by amplicon sequencing (NGS) of the target-to-LE donor junction, after delivery via lipofection of a plasmid containing all CAST protein components, a donor plasmid containing a full factor IX gene and the targeting sgRNA.Attorney Docket No. MTGNZ000100WO

[0019] FIG. 9 illustrates how MG64- 1 TnsB variants from combinatorial libraries improves integration efficiency >4%. Integration efficiency in HEK293T cells with MG64-1 TnsB engineered variants based on combinatorial mutations of up to 10 selected sites. Integration efficiency was measured by amplicon sequencing (NGS) of the target-to-LE donor junction, after delivery via lipofection of a plasmid containing all CAST protein components, a donor plasmid containing a full factor IX geneand the targeting sgRNA, and a plasmid containing ClpX. TnsB hexa: SEQ ID NO: 1430, TnsB tetra: SEQ ID NO: 1431.BRIEF DESCRIPTION OF THE SEQUENCE LISTING

[0020] The Sequence Listing filed herewith provides exemplary polynucleotide and polypeptide sequences for use in methods, compositions, and systems according to the disclosure. Below are exemplary descriptions of sequences therein.

[0021] MG64

[0022] SEQ ID NOs: 1 , 12, 16, 20-30, 64, 80-85, and 220 show the full-length peptide sequences of MG64 Cas effectors.

[0023] SEQ ID NOs: 2-4, 13-15, 17-19, 65-67, and 109-111 show the peptide sequences of MG64 transposition proteins that may comprise a recombinase / transposase recognition complex associated with the MG64 Cas effector.

[0024] SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, 119-122, and 222 show nucleotide sequences of MG64 tracrRNAs derived from the same loci as a MG64 Cas effector.

[0025] SEQ ID NOs: 7 and 34-35 show nucleotide sequences of MG64 target CRISPR repeats.

[0026] SEQ ID NOs: 106-108, 112-118, and 221 show nucleotide sequences of MG64 crRNAs.

[0027] SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93 show nucleotide sequences of right-hand transposase recognition sequences associated with a MG64 system.

[0028] SEQ ID NOs: 9, 11, 36-38, 76, and 78 show nucleotide sequences of left-hand transposase recognition sequences associated with a MG64 system.

[0029] SEQ ID NOs: 45-63, 68-75, 96-103, and 123-140 show nucleotide sequences of single guide RNAs engineered to function with MG64 Cas effectors.

[0030] SEQ ID NO: 208 shows the nucleotide sequence of an MG64 expression construct.

[0031] SEQ ID NO: 223 shows the nucleotide sequence of an MG64 active donor.

[0032] SEQ ID NOs: 228-230 show the full-length peptide sequences of MG64 accessory proteins.

[0033] SEQ ID NOs: 233-234 show the nucleotide sequences of MG64 target sites.Attorney Docket No. MTGNZ000100WO

[0034] SEQ ID NOs: 369-371 show the nucleotide sequences of MG64 active target sequences.

[0035] MG190

[0036] SEQ ID NOs: 209-219 show the full-length peptide sequences of MG190 ribosomal protein S15 homologs.

[0037] Other Sequences

[0038] SEQ ID NOs: 86-87,192-207, and 1354-1383 show peptide sequences of nuclear localizing signals.

[0039] SEQ ID NOs: 88-89 show peptide sequences of linkers.

[0040] SEQ ID NOs: 90-92 show peptide sequences of epitope tags.

[0041] SEQ ID NOs: 141-143 show genomic target sequences.

[0042] SEQ ID NOs: 144-180 show target guide sequences.

[0043] SEQ ID NOs: 181-183 show nucleic acid sequences of the S 15 fusion proteins.

[0044] SEQ ID NO: 184 shows a donor construct.

[0045] SEQ ID NO: 185 shows an MG64-1 sgRNA sequence.

[0046] SEQ ID NO: 186 shows a linker sequence.

[0047] SEQ ID NOs: 187-189 show amino acid sequences of the S 15 fusion proteins.

[0048] SEQ ID NOs: 190-191 show promoter sequences.

[0049] SEQ ID NOs: 224-226, 231-232, 250-251, 253, 1385, 1386, 1388, 1389, and 1391 show the nucleotide sequences of primers.

[0050] SEQ ID NO: 227 shows the nucleotide sequence of a plasmid element.

[0051] SEQ ID NOs: 235-249 show the peptide sequences of ClpX accessory proteins.

[0052] SEQ ID NO: 252 shows the nucleotide sequence of a plasmid element.

[0053] SEQ ID NOs: 254-256 show the nucleotide sequences of MG64 active target sequences.

[0054] SEQ ID NOs: 257-282 and 1138-1241 show the protein sequences of MG161 functional domains.

[0055] SEQ ID NOs: 283-307 and 1242-1254 show the protein sequences of MG162 functional domains.

[0056] SEQ ID NOs: 308-359 show the protein sequences of Hl core library.

[0057] SEQ ID NOs: 360-368 and 372-753 show the nucleotide sequences of primers and probes.

[0058] SEQ ID NOs: 754-944 show the nucleotide sequences of single guide targets.

[0059] SEQ ID NOs: 945-1135 show the nucleotide sequences of primer binding sequences.

[0060] SEQ ID NO: 1136 shows a nucleotide sequence of a promoter.Attorney Docket No. MTGNZ000100WO

[0061] SEQ ID NO: 1137 shows a protein sequence of an expression construct.

[0062] SEQ ID NO: 1387 shows a nucleotide sequence of an AAVS1 target sequence.

[0063] SEQ ID NO: 1390 shows a nucleotide sequence of AAVS1 target pdonor sequence.

[0064] SEQ ID NO: 1392 shows a nucleotide sequence of a transposed sequence.

[0065] SEQ ID NO: 1393 shows a nucleotide sequence of an untransposed sequence.

[0066] SEQ ID NO: 1394 shows a nucleotide sequence of S15-NLS.

[0067] SEQ ID NO: 1395 shows a nucleotide sequence of ClpX-NLS.

[0068] SEQ ID NO: 1396 shows a nucleotide sequence of a terminal inverted repeat (TIR)- left-end recognition sequence of a .S', hoffmanni type V-K CAST (ShCAST) system described herein.

[0069] SEQ ID NO: 1397 shows a nucleotide sequence of a TIR-right-end recognition sequence of a ShCAST system described herein.

[0070] SEQ ID NOs: 1398-1401 show nucleotide sequences of ShCAST systems.

[0071] SEQ ID NO: 1402 shows a nucleotide sequence of an AAVS1 sgRNA for use with the ShCAST systems described herein.

[0072] SEQ ID NO: 1403 shows a nucleotide sequence of a null sgRNA for use with the ShCAST systems described herein.

[0073] SEQ ID NO: 1404 shows a nucleotide sequence of a TIR-left-end recognition sequence of a MG64-1 system described herein.

[0074] SEQ ID NO: 1405 shows a nucleotide sequence of a TIR-right-end recognition sequence of a MG64-1 system described herein.

[0075] SEQ ID NOs: 1406-1409 show nucleotide sequences of MG64-1 CAST systems.

[0076] SEQ ID NO: 1410 shows a nucleotide sequence of an AAVS1 sgRNA for use with the MG64-1 systems described herein.

[0077] SEQ ID NO: 1411 shows a nucleotide sequence of a null sgRNA for use with the MG64-1 systems described herein.

[0078] SEQ ID NOs: 1412-1414 show primers used with MG64-1. SEQ ID NOs: 1415 and 1416 show MG64-1 active target sequences. SEQ ID NOs: 1417 shows an expression construct. SEQ ID NO: 1424 shows a single guide target. SEQ ID NOs: 1418-1424 show MG64-1-B variants.DETAILED DESCRIPTION

[0079] While various embodiments of the disclosure have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled inAttorney Docket No. MTGNZ000100WO the art without departing from the disclosure. It should be understood that various alternatives to the embodiments of the disclosure described herein may be employed.

[0080] The practice of some methods disclosed herein employ, unless otherwise indicated, techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA. See for example Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (F. M. Ausubel, et al. eds.); the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (M.J. MacPherson, B.D. Hames and G.R. Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (R.I. Freshney, ed. (2010)).

[0081] As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, to the extent that the terms “including”, “includes”, “having”, “has”, “with”, or variants thereof are used in either the detailed description and / or the claims, such terms are intended to be inclusive in a manner similar to the term “comprising.”

[0082] The term “about” or “approximately” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, “about” can mean within one or more than one standard deviation, per the practice in the art. Alternatively, “about” can mean a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1 % of a given value.

[0083] The term “nucleotide,” as used herein, refers to a base-sugar-phosphate combination. Contemplated nucleotides include naturally occurring nucleotides and synthetic nucleotides. Nucleotides are monomeric units of a nucleic acid sequence (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide includes ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP) and deoxyribonucleoside triphosphates such as dATP, dCTP, diTP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives include, for example, [aS]dATP, 7-deaza-dGTP and 7-deaza-dATP, and nucleotide derivatives that confer nuclease resistance on the nucleic acid molecule containing them. The term nucleotide as used herein encompasses dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Illustrative examples of ddNTPs include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. A nucleotide may be unlabeled or detectably labeled, such as using moieties comprising optically detectable moieties (e.g., fluorophores) or quantum dots.Attorney Docket No. MTGNZOOOIOOWODetectable labels include, for example, radioactive isotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels. Fluorescent labels of nucleotides include but are not limited fluorescein, 5-carboxyfluorescein (FAM), 2'7'- dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4- (4'dimethylaminophenylazo) benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, Cyanine and 5-(2'-aminoethyl)aminonaphthalene-l -sulfonic acid (EDANS). Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [RU0]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP available from Perkin Elmer, Foster City, Calif; FluoroLink Deoxy Nucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP available from Amersham, Arlington Heights, IL; Fluorescein- 15-dATP, Fluorescein- 12-dUTP, Tetramethyl-rodamine-6- dUTP, IR770-9-dATP, Fluorescein- 12-ddUTP, Fluorescein- 12-UTP, and Fluorescein- 15-2'- dATP available from Boehringer Mannheim, Indianapolis, Ind.; and Chromosome Labeled Nucleotides, BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, B0DIPY-TMR-14-UTP, BODIPY- TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, fluorescein- 12-UTP, fluorescein- 12-dUTP, Oregon Green 488-5- dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, tetramethylrhodamine-6-UTP, tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red- 12- dUTP available from Molecular Probes, Eugene, Oreg. The term nucleotide encompasses chemically modified nucleotides. An exemplary chemically-modified nucleotide is biotin- dNTP. Non-limiting examples of biotinylated dNTPs include, biotin-dATP (e.g., bio-N6- ddATP, biotin- 14-dATP), biotin-dCTP (e.g., biotin- 11-dCTP, biotin- 14-dCTP), and biotin- dUTP (e.g., biotin- 11 -dUTP, biotin- 16-dUTP, biotin-20-dUTP).

[0084] The terms “polynucleotide,” “oligonucleotide,” and “nucleic acid” are used interchangeably to refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof, either in single-, double-, or multistranded form. Contemplated polynucleotides include a gene or fragment thereof. Exemplary polynucleotides include, but are not limited to, DNA, RNA, coding or non-coding regions of a gene or gene fragment, loci (locus) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short-hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, cell-free polynucleotidesAttorney Docket No. MTGNZOOOIOOWO including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. In a polynucleotide when referring to a T, a T means U (Uracil) in RNA and T (Thymine) in DNA. A polynucleotide can be exogenous or endogenous to a cell and / or exist in a cell-free environment. The term polynucleotide encompasses modified polynucleotides (e.g., altered backbone, sugar, or nucleobase). If present, modifications to the nucleotide structure are imparted before or after assembly of the polymer. Non-limiting examples of modifications include: 5-bromouracil, peptide nucleic acid, xeno nucleic acid, morpholines, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxy nucleotides, cordycepin, 7-deaza- GTP, fluorophores (e.g., rhodamine or fluorescein linked to the sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7- guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queuosine, and wyosine. The sequence of nucleotides may be interrupted by non-nucleotide components.

[0085] The terms “peptide,” “polypeptide,” and “protein” are used interchangeably herein to refer to a polymer of at least two amino acid residues joined by peptide bond(s). This term does not connote a specific length of polymer, nor is it intended to imply or distinguish whether the peptide is produced using recombinant techniques, chemical or enzymatic synthesis, or is naturally occurring. The terms apply to naturally occurring amino acid polymers as well as amino acid polymers comprising at least one modified amino acid. In some cases, the polymer is interrupted by non-amino acids. The terms include amino acid chains of any length, including full length proteins, and proteins with or without secondary or tertiary structure (e.g., domains). The terms also encompass an amino acid polymer that has been modified, for example, by disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, oxidation, and any other manipulation such as conjugation with a labeling component. The terms “amino acid” and “amino acids,” as used herein, refer to natural and non-natural amino acids, including, but not limited to, modified amino acids. Modified amino acids include amino acids that have been chemically modified to include a group or a chemical moiety not naturally present on the amino acid. The term “amino acid” includes both D-amino acids and L- amino acids.

[0086] As used herein, the “non-native” refers to a nucleic acid or polypeptide sequence that is non-naturally occurring. Non-native refers to a non-naturally occurring nucleic acid or polypeptide sequence that comprises modifications such as mutations, insertions, or deletions. The term non-native encompasses fusion nucleic acids or polypeptides that encodes or exhibits an activity (e.g., enzymatic activity, methyltransferase activity, acetyltransferase activity, kinase activity, ubiquitinating activity, etc.) of the nucleic acid or polypeptide sequence toAttorney Docket No. MTGNZ000100WO which the non-native sequence is fused. A non-native nucleic acid or polypeptide sequence includes those linked to a naturally-occurring nucleic acid or polypeptide sequence (or a variant thereof) by genetic engineering to generate a chimeric nucleic acid or polypeptide sequence encoding a chimeric nucleic acid or polypeptide.

[0087] As used herein, “operably linked”, “operable linkage”, “operatively linked”, or grammatical equivalents thereof refer to an arrangement of genetic elements, e.g., a promoter, an enhancer, a polyadenylation sequence, etc., wherein an operation (e.g., movement or activation) of a first genetic element has some effect on the second genetic element. The effect on the second genetic element can be, but need not be, of the same type as operation of the first genetic element. For example, two genetic elements are operably linked if movement of the first element causes an activation of the second element. For instance, a regulatory element, which may comprise promoter and / or enhancer sequences, is operatively linked to a coding region if the regulatory element helps initiate transcription of the coding sequence. There may be intervening residues between the regulatory element and coding region so long as this functional relationship is maintained.

[0088] A “functional fragment” of a DNA or protein sequence refers to a fragment that retains a biological activity (either functional or structural) that is substantially similar to a biological activity of the full-length DNA or protein sequence. A biological activity of a DNA sequence includes its ability to influence expression in a manner attributed to the full-length sequence.

[0089] The terms “engineered,” “synthetic,” and “artificial” are used interchangeably herein to refer to an object that has been modified by human intervention. For example, the terms refer to a polynucleotide or polypeptide that is non- naturally occurring. An engineered peptide has, but does not require, low sequence identity (e.g., less than 50% sequence identity, less than 25% sequence identity, less than 10% sequence identity, less than 5% sequence identity, less than 1 % sequence identity) to a naturally occurring human protein. For example, VPR and VP64 domains are synthetic transactivation domains. Non-limiting examples include the following: a nucleic acid modified by changing its sequence to a sequence that does not occur in nature; a nucleic acid modified by ligating it to a nucleic acid that it does not associate with in nature such that the ligated product possesses a function not present in the original nucleic acid; an engineered nucleic acid synthesized in vitro with a sequence that does not exist in nature; a protein modified by changing its amino acid sequence to a sequence that does not exist in nature; an engineered protein acquiring a new function or property. An “engineered” system comprises at least one engineered component.

[0090] The term “tracrRNA” or “tracr sequence” means trans-activating CRISPR RNA. tracrRNA interacts with the CRISPR (cr) RNA to form a guide nucleic acid (e.g., guide RNAAttorney Docket No. MTGNZ000100WO or gRNA) that may hybridize to a target nucleic acid and thereby directs an associated nuclease to the target nucleic acid.

[0091] As used herein, a “guide nucleic acid” or “guide polynucleotide” refers to a nucleic acid that may hybridize to a target nucleic acid and thereby directs an associated nuclease to the target nucleic acid. A guide nucleic acid is, but is not limited to, RNA (guide RNA or gRNA), DNA, or a mixture of RNA and DNA. A guide nucleic acid can include a crRNA or a tracrRNA or a combination of both. The term guide nucleic acid encompasses an engineered guide nucleic acid and a programmable guide nucleic acid to specifically bind to the target nucleic acid. A portion of the target nucleic acid may be complementary to a portion of the guide nucleic acid. The strand of a double-stranded target polynucleotide that is complementary to and hybridizes with the guide nucleic acid is the complementary strand. The strand of the double-stranded target polynucleotide that is complementary to the complementary strand, and therefore is not complementary to the guide nucleic acid is called noncomplementary strand. A guide nucleic acid having a polynucleotide chain is a “single guide nucleic acid.” A guide nucleic acid having two polynucleotide chains is a “double guide nucleic acid.” If not otherwise specified, the term “guide nucleic acid” is inclusive, referring to both single guide nucleic acids and double guide nucleic acids. A guide nucleic acid may comprise a segment referred to as a “nucleic acid-targeting segment” or a “nucleic acidtargeting sequence,” or a “spacer.” A nucleic acid- targeting segment can include a sub-segment referred to as a “protein binding segment” or “protein binding sequence” or “Cas protein binding segment.”

[0092] The term “sequence identity” or “percent identity” in the context of two or more nucleic acids or polypeptide sequences, refers to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences that are the same or have a specified percentage of amino acid residues or nucleotides that are the same, when compared and aligned for maximum correspondence over a local or global comparison window, as measured using a sequence comparison algorithm. Suitable sequence comparison algorithms for polypeptide sequences include, e.g., BLASTP using parameters of a wordlength (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix setting gap costs at existence of 11, extension of 1, and using a conditional compositional score matrix adjustment for polypeptide sequences longer than 30 residues; BLASTP using parameters of a wordlength (W) of 2, an expectation (E) of 1000000, and the PAM30 scoring matrix setting gap costs at 9 to open gaps and 1 to extend gaps for sequences of less than 30 residues (these are the default parameters for BLASTP in the BLAST suite available at https: / / blast.ncbi.nlm.nih.gov); CLUSTALW with the Smith-Waterman homology search algorithm parameters with a match of 2, a mismatch ofAttorney Docket No. MTGNZ000100WO-1, and a gap of -1 ; MUSCLE with default parameters; MAFFT with parameters of a retree of 2 and max iterations of 1000; Novafold with default parameters; HMMER hmmalign with default parameters.

[0093] Included in the current disclosure are variants of any of the enzymes described herein with one or more conservative amino acid substitutions. Such conservative substitutions can be made in the amino acid sequence of a polypeptide without disrupting the three-dimensional structure or function of the polypeptide. Conservative substitutions can be accomplished by substituting amino acids with similar hydrophobicity, polarity, and R chain length for one another. Additionally or alternatively, by comparing aligned sequences of homologous proteins from different species, conservative substitutions can be identified by locating amino acid residues that have been mutated between species (e.g., non-conserved residues without altering the basic functions of the encoded proteins. Such conservatively substituted variants may include variants with at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity any one of the systems described herein (e.g., MG64 systems described herein). In some embodiments, such conservatively substituted variants are functional variants. Such functional variants can encompass sequences with substitutions such that the activity of critical active site residues of the endonuclease are not disrupted.

[0094] Also included in the current disclosure are variants of any of the enzymes described herein with substitution of one or more catalytic residues to decrease or eliminate activity of the enzyme (e.g. decreased-activity variants). In some embodiments, a decreased activity variant as a protein described herein comprises a disrupting substitution of at least one, at least two, or all three catalytic residues.

[0095] Conservative substitution tables providing functionally similar amino acids are available from a variety of references (see, for example, Creighton, Proteins: Structures and Molecular Properties (W H Freeman & Co.; 2ndEdition (December 1993))). The following eight groups each contain amino acids that are conservative substitutions for one another:1) Alanine (A), Glycine (G);2) Aspartic acid (D), Glutamic acid (E);3) Asparagine (N), Glutamine (Q);4) Arginine (R), Lysine (K);5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V);Attorney Docket No. MTGNZ000100WO6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W);7) Serine (S), Threonine (T); and8) Cysteine (C), Methionine (M).

[0096] As used herein, the term “RuvC_III domain” refers to a third discontinuous segment of a RuvC endonuclease domain (the RuvC nuclease domain being comprised of three discontiguous segments, RuvC_I, RuvC_II, and RuvC_III). A RuvC domain or segments thereof can generally be identified by alignment to documented domain sequences, structural alignment to proteins with annotated domains, or by comparison to Hidden Markov Models (HMMs) built based on documented domain sequences (e.g., Pfam HMM PF18541 for RuvCJII).

[0097] As used herein, the term “HNH domain” refers to an endonuclease domain having characteristic histidine and asparagine residues. An HNH domain can generally be identified by alignment to documented domain sequences, structural alignment to proteins with annotated domains, or by comparison to Hidden Markov Models (HMMs) built based on documented domain sequences (e.g., Pfam HMM PF01844 for domain HNH).

[0098] As used herein, the term “recombinase” refers to an enzyme that mediates the recombination of DNA fragments located between recombinase recognition sequences, which results in the excision, insertion, inversion, exchange or translocation) of the DNA fragments located between the recombinase recognition sequences.

[0099] As used herein, the term “recombine,” or “recombination,” in the context of a nucleic acid modification (e.g., a genomic modification), refers to the process by which two or more nucleic acid molecules, or two or more regions of a single nucleic acid molecule, are modified by the action of a recombinase protein. Recombination can result in, inter alia, the excision, insertion, inversion, exchange, or translocation of a nucleic acid sequence, e.g., in or between one or more nucleic acid molecules.

[0100] As used herein, the term “transposon,” or “transposable element” refers to a nucleic acid sequence in a genome that is a mobile genetic element that can change its position in a genome. In some embodiments, the transposon transports additional “cargo DNA” excised from the genome. Transposons comprise, for example retrotransposons, DNA transposons, autonomous and non-autonomous transposons, and class III transposons. Transposon nucleic acid sequences comprise, for example genes coding for a cognate transposase, one or more recognition sequences for the transposase, or combinations thereof. In some embodiments, these transposons differ on the type of nucleic acid to transpose, the type of repeat at the ends of the transposon, the type of cargo to be carried or by the mode of transposition (i.e. selfrepair or host-repair). As used herein, the term “transposase” or “transposases” refers to anAttorney Docket No. MTGNZOOOIOOWO enzyme that binds to the recognition sequences of a transposon and catalyzes its movement to another part of the genome. In some embodiments, the movement is by a cut and paste mechanism or a replicative transposition mechanism.

[0101] As used herein, the term “Tn7” or “Tn7-like transposase” refers to a family of transposases comprising three main components: a heteromeric transposase (TnsA and / or TnsB) alongside a regulator protein (TnsC). In addition to the TnsABC transposition proteins, Tn7 elements can encode dedicated target site-selection proteins, TnsD and TnsE. In conjunction with TnsABC, the sequence-specific DNA-binding protein TnsD directs transposition into a conserved site referred to as the “Tn7 attachment site,” attTn7. TnsD is a member of a large family of proteins that also includes TniQ. TniQ has been shown to target transposition into resolution sites of plasmids.

[0102] As used herein, the term “complex” refers to a joining of at least two components. The two components may each retain the properties / activities they had prior to forming the complex. The joining may be by covalent bonding, non-covalent bonding (i.e., hydrogen bonding, ionic interactions, Van der Waals interactions, and hydrophobic bond), use of a linker, fusion, or any other suitable method. In some embodiments, components in a complex are polynucleotides, polypeptides, or combinations thereof. For example, a complex may comprise a Cas protein and a guide nucleic acid.

[0103] In some embodiments, the CAST systems described herein comprise one or more Tn7 or Tn7 like transposases. In certain example embodiments, the Tn7 or Tn7 like transposase comprises a multimeric protein complex. In certain example embodiments, the multimeric protein complex comprises TnsA, TnsB, TnsC, or TniQ. In these combinations, the transposases (TnsA, TnsB, TnsC, TniQ) may form complexes or fusion proteins with each other.

[0104] In some embodiments, the CAST systems described herein comprise one or more Tn5O53 or Tn5053-like transposases. In certain example embodiments, the Tn5O53 or Tn5053- like transposase comprises a multimeric protein complex. In certain example embodiments, the multimeric protein complex comprises TnsA, TnsB, TnsC, or TniQ. In these combinations, the transposases (TnsA, TnsB, TnsC, TniQ) may form complexes or fusion proteins with each other.

[0105] As used herein, the term “Cas 12k” (alternatively “class 2, type V-K”) refers to a subtype of Type V CRISPR systems that have been found to be defective in nuclease activity (e.g., they may comprise at least one defective RuvC domain that lacking at least one catalytic residue important for DNA cleavage). Such subtype of effectors have been generally associated with CAST systems.Attorney Docket No. MTGNZOOOIOOWO

[0106] In accordance with IUPAC conventions, the following abbreviations are used throughout the examples: A = adenine C = cytosine G = guanine T = thymine R = adenine or guanineY = cytosine or thymine S - guanine or cytosine W = adenine or thymine K = guanine or thymine M = adenine or cytosine B = C, G, or TD = A, G, or T H = A, C, or TV = A, C, or GOverview

[0107] The discovery of new Cas enzymes with unique functionality and structure may offer the potential to further disrupt deoxyribonucleic acid (DNA) editing technologies, improving speed, specificity, functionality, and ease of use. Relative to the predicted prevalence of Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) systems in microbes and the sheer diversity of microbial species, relatively few functionally characterized CRISPR / Cas enzymes exist in the literature. This is partly because a huge number of microbial species may not be readily cultivated in laboratory conditions. Metagenomic sequencing from natural environmental niches that represent large numbers of microbial species may offer the potential to drastically increase the number of new CRISPR / Cas systems documented and speed the discovery of new oligonucleotide editing functionalities. A recent example of the fruitfulness of such an approach is demonstrated by the 2016 discovery of CasX / CasY CRISPR systems from metagenomic analysis of natural microbial communities.

[0108] CRISPR / Cas systems are RNA-directed nuclease complexes that have been described to function as an adaptive immune system in microbes. In their natural context, CRISPR / Cas systems occur in CRISPR (clustered regularly interspaced short palindromic repeats) operons or loci, which generally comprise two parts: (i) an array of short repetitive sequences (30-40 bp) separated by equally short spacer sequences, which encode the RNA-based targeting element; and (ii) ORFs encoding the Cas encoding the nuclease polypeptide directed by the RNA-based targeting element alongside accessory proteins / enzymes. Efficient nuclease targeting of a particular target nucleic acid sequence generally requires both (i) complementary hybridization between the first 6-8 nucleic acids of the target (the target seed) and the crRNAAttorney Docket No. MTGNZOOOIOOWO guide; and (ii) the presence of a protospacer-adjacent motif (PAM) sequence within a defined vicinity of the target seed (the PAM usually being a sequence not commonly represented within the host genome). Depending on the exact function and organization of the system, CRISPR-Cas systems are commonly organized into 2 classes, 5 types and 16 subtypes based on shared functional characteristics and evolutionary similarity.

[0109] Class 1 CRISPR-Cas systems have large, multisubunit effector complexes, and comprise Types I, III, and IV.

[0110] Type I CRISPR-Cas systems are considered of moderate complexity in terms of components. In Type I CRISPR-Cas systems, the array of RNA-targeting elements is transcribed as a long precursor crRNA (pre-crRNA) that is processed at repeat elements to liberate short, mature crRNAs that direct the nuclease complex to nucleic acid targets when they are followed by a suitable short consensus sequence called a protospacer- adjacent motif (PAM). This processing occurs via an endoribonuclease subunit (Cas6) of a large endonuclease complex called Cascade, which also comprises a nuclease (Cas3) protein component of the crRNA-directed nuclease complex. Cas I nucleases function primarily as DNA nucleases.

[0111] Type III CRISPR systems may be characterized by the presence of a central nuclease, known as Cas 10, alongside a repeat-associated mysterious protein (RAMP) that comprises Csm or Cmr protein subunits. Like in Type I systems, the mature crRNA is processed from a pre-crRNA using a Cas6-like enzyme. Unlike type I and II systems, type III systems appear to target and cleave DNA-RNA duplexes (such as DNA strands being used as templates for an RNA polymerase).

[0112] Type IV CRISPR-Cas systems possess an effector complex that comprises a highly reduced large subunit nuclease (csfl ), two genes for RAMP proteins of the Cas5 (csf3) and Cas7 (csf2) groups, and, in some embodiments, a gene for a predicted small subunit; such systems are commonly found on endogenous plasmids.

[0113] Class 2 CRISPR-Cas systems generally have single-polypeptide multidomain nuclease effectors, and comprise Types II, V and VI.

[0114] Type II CRISPR-Cas systems are considered the simplest in terms of components. In Type II CRISPR-Cas systems, the processing of the CRISPR array into mature crRNAs does not require the presence of a special endonuclease subunit, but rather a small trans-encoded crRNA (tracrRNA) with a region complementary to the array repeat sequence; the tracrRNA interacts with both its corresponding effector nuclease (e.g., Cas9) and the repeat sequence to form a precursor dsRNA structure, which is cleaved by endogenous RNAse III to generate a mature effector enzyme loaded with both tracrRNA and crRNA. Type II nucleases are known as DNA nucleases. Type II effectors generally exhibit a structure consisting of a RuvC-likeAttorney Docket No. MTGNZOOOIOOWO endonuclease domain that adopts the RNase H fold with an unrelated HNH nuclease domain inserted within the folds of the RuvC-like nuclease domain. The RuvC-like domain is responsible for the cleavage of the target (e.g., crRNA complementary) DNA strand, while the HNH domain is responsible for cleavage of the displaced DNA strand.

[0115] Type V CRISPR-Cas systems are characterized by a nuclease effector (e.g., Casl2) structure similar to that of Type II effectors, comprising a RuvC-like domain. Similar to Type II, most (but not all) Type V CRISPR systems use a tracrRNA to process pre-crRNAs into mature crRNAs; however, unlike Type II systems which requires RNAse III to cleave the pre- crRNA into multiple crRNAs, Type V systems are capable of using the effector nuclease itself to cleave pre-crRNAs. Like Type-II CRISPR-Cas systems, Type V CRISPR-Cas systems are again known as DNA nucleases. Unlike Type II CRISPR-Cas systems, some Type V enzymes (e.g., Casl2a) appear to have a robust single- stranded nonspecific deoxyribonuclease activity that is activated by the first crRNA directed cleavage of a double-stranded target sequence.

[0116] Type VI CRISPR-Cas systems have RNA-guided RNA endonucleases. Instead of RuvC-like domains, the single polypeptide effector of Type VI systems (e.g., Casl3) comprises two HEPN ribonuclease domains. Differing from both Type II and V systems, Type VI systems also appear to not need a tracrRNA for processing of pre-crRNA into crRNA. Similar to type V systems, however, some Type VI systems (e.g., C2C2) appear to possess robust single- stranded nonspecific nuclease (ribonuclease) activity activated by the first crRNA directed cleavage of a target RNA.

[0117] Because of their simpler architecture, Class 2 CRISPR-Cas have been most widely adopted for engineering and development as designer nuclease / genome editing applications.

[0118] One of the early adaptations of such a system for in vitro use involved (i) recombinantly-expressed, purified full-length Cas9 (e.g., a Class 2, Type II Cas enzyme) isolated from S. pyogenes SF370, (ii) purified mature ~42 nt crRNA bearing a ~20 nt 5’ sequence complementary to the target DNA sequence desired to be cleaved followed by a 3’ tracr-binding sequence (the whole crRNA being in vitro transcribed from a synthetic DNA template carrying a T7 promoter sequence); (iii) purified tracrRNA in vitro transcribed from a synthetic DNA template carrying a T7 promoter sequence, and (iv) Mg2+. A later improved, engineered system involved the crRNA of (ii) joined to the 5’ end of (iii) by a linker (e.g., GAAA) to form a single fused synthetic guide RNA (sgRNA) capable of directing Cas9 to a target by itself.

[0119] Such engineered systems can be adapted for use in mammalian cells by providing DNA vectors encoding (i) an ORF encoding codon-optimized Cas9 (e.g., a Class 2, Type II Cas enzyme) under a suitable mammalian promoter with a C-terminal nuclear localizationAttorney Docket No. MTGNZOOOIOOWO sequence (e.g., SV40 NLS) and a suitable polyadenylation signal (e.g., TK pA signal); and (ii) an ORF encoding an sgRNA (having a 5 ’ sequence beginning with G followed by 20 nt of a complementary targeting nucleic acid sequence joined to a 3’ tracr-binding sequence, a linker, and the tracrRNA sequence) under a suitable Polymerase III promoter (e.g., the U6 promoter).

[0120] Transposons are mobile elements that can move between positions in a genome. Such transposons have evolved to limit the negative effects they exert on the host. A variety of regulatory mechanisms are used to maintain transposition at a low frequency and sometimes coordinate transposition with various cell processes. Some prokaryotic transposons also can mobilize functions that benefit the host or otherwise help maintain the element. Certain transposons may have also evolved mechanisms of tight control over target site selection, the most notable example being the Tn7 family.

[0121] Transposon Tn7 and similar elements may be reservoirs for antibiotic resistance and pathogenesis functions in clinical settings, as well as encoding other adaptive functions in natural environments. The Tn7 system, for example, has evolved mechanisms to almost completely avoid integrating into important host genes, but also maximize dispersal of the element by recognizing mobile plasmids and bacteriophage capable of moving Tn7 between host bacteria.

[0122] Tn7 and Tn7-like elements may control where and when they insert, possessing one pathway that directs insertion into a single conserved position in bacterial genomes and a second pathway that appears to be adapted to maximizing targeting into mobile plasmids capable of transporting the element between bacteria. The association between Tn7-like transposons and CRISPR-Cas systems suggests that the transposons might have hijacked CRISPR effectors to generate R-loops in target sites and facilitate the spread of transposons via plasmids and phages.MG64 Systems

[0123] Provided herein, in some embodiments, are MG64 systems for transposing a cargo nucleotide sequence into a target nucleic acid site.

[0124] Described herein, in certain embodiments, are systems for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid comprising: a) a Cas effector complex and a Tn7 type transposase complex that binds the Cas effector complex, wherein the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence having at least 70% sequence identity to any one of SEQ ID NOs: 1406- 1409; and b) a double-stranded nucleic acid configured to interact with the Tn7 type transposase complex and comprising the cargo nucleotide sequence.Attorney Docket No. MTGNZ000100WO

[0125] Further described herein, in certain embodiments, are systems for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid comprising: a) a Cas effector complex and a Tn7 type transposase complex that binds the Cas effector complex, wherein the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence having at least 70% sequence identity to any one of SEQ ID NOs: 1406- 1409; and b) a double-stranded nucleic acid that interacts with the Tn7 type transposase complex and comprising in 5’ to 3’ order: i) a left-hand transposase recognition sequence comprising a nucleotide sequence having at least 80% sequence identity to SEQ ID NO: 1404; ii) the cargo nucleotide sequence; and iii) a right-hand transposase recognition sequence comprising a nucleotide sequence having at least 80% identity to SEQ ID NO: 1405.

[0126] In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 70% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 75% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 80% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 85% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 90% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 91% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 92% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 93% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complexAttorney Docket No. MTGNZ000100WO comprise a nucleotide sequence having at least about 94% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 95% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 96% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 97% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 98% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 99% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having 100% identity to any one of SEQ ID NOs: 1406-1409.

[0127] In some embodiments, the Cas effector complex comprises a class 2, type V Cas effector, a small prokaryotic ribosomal protein subunit SI 5, an engineered guide polynucleotide configured to hybridize to the target nucleic acid site, or combinations thereof.

[0128] In some embodiments, the Tn7 type transposase complex comprises a TnsB, TnsC, and TniQ component and / or an accessory protein. In some embodiments, the Tn7 type transposase complex comprises a functional domain (FD)-TniQ fusion and / or an accessory protein.

[0129] In some embodiments, this cargo nucleotide sequence interacts with a Tn7 type or Tn5O53 type transposase complex. In some embodiments, the system comprises a Tn7 type or Tn5O53 type transposase complex configured to bind the Cas effector complex, wherein the Tn7 type or Tn5053 type transposase complex comprises a TnsB subunit.

[0130] In some embodiments, the class 2, type V Cas effector and the Tn7 type transposase complex are encoded by polynucleotide sequences comprising fewer than about 20 kilobases, fewer than about 15 kilobases, fewer than about 10 kilobases, or fewer than about 5 kilobases.

[0131] In some embodiments, the Cas effector complex comprises a class 2, type V Cas effector. In some embodiments, the class 2, type V Cas effector is a class 2, type V-K effector. In some embodiments, the class 2, type V Cas effector comprises a polypeptide comprising a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%,Attorney Docket No. MTGNZ000100WO or at least about 99% identity to SEQ ID NO: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the class 2, type V Cas effector comprises a polypeptide comprising a sequence having at least about 70% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the class 2, type V Cas effector comprises a polypeptide comprising a sequence having at least about 75% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the class 2, type V Cas effector comprises a polypeptide comprising a sequence having at least about 80% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the class 2, type V Cas effector comprises a polypeptide comprising a sequence having at least about 85% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the class 2, type V Cas effector comprises a polypeptide comprising a sequence having at least about 90% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the class 2, type V Cas effector comprises a polypeptide comprising a sequence having at least about 91% identity to SEQ ID NOs: 1 , 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the class 2, type V Cas effector comprises a polypeptide comprising a sequence having at least about 92% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the class 2, type V Cas effector comprises a polypeptide comprising a sequence having at least about 93% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the class 2, type V Cas effector comprises a polypeptide comprising a sequence having at least about 94% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the class 2, type V Cas effector comprises a polypeptide comprising a sequence having at least about 95% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the class 2, type V Cas effector comprises a polypeptide comprising a sequence having at least about 96% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the class 2, type V Cas effector comprises a polypeptide comprising a sequence having at least about 97% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the class 2, type V Cas effector comprises a polypeptide comprising a sequence having at least about 98% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the class 2, type V Cas effector comprises a polypeptide comprising a sequence having at least about 99% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220. In some embodiments, the class 2, type V Cas effector comprises a polypeptide comprising a sequence having 100% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220.

[0132] In some embodiments, the Tn7 type transposase complex comprises a TnsB subunit. In some embodiments, the TnsB subunit comprises a polypeptide having a sequence having atAttorney Docket No. MTGNZ000100WO least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 2, 13, 17, and 65. In some embodiments, the TnsB subunit comprises a polypeptide having a sequence identical to SEQ ID NO: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide comprising a sequence having at least about 70% identity to SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide comprising a sequence having at least about 75% identity to SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide comprising a sequence having at least about 80% identity to SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide comprising a sequence having at least about 85% identity to SEQ ID NOs: 2, 13, 17, and 65. Tn some embodiments, the TnsB component comprises a polypeptide comprising a sequence having at least about 90% identity to SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide comprising a sequence having at least about 91% identity to SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide comprising a sequence having at least about 92% identity to SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide comprising a sequence having at least about 93% identity to SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide comprising a sequence having at least about 94% identity to SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide comprising a sequence having at least about 95% identity to SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide comprising a sequence having at least about 96% identity to SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide comprising a sequence having at least about 97% identity to SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide comprising a sequence having at least about 98% identity to SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide comprising a sequence having at least about 99% identity to SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide comprising a sequence having 100% identity to SEQ ID NOs: 2, 13, 17, and 65.

[0133] In some embodiments, the Tn7 type transposase complex comprises a functional domain (FD)-TniQ fusion. In some embodiments, the functional domain (FD)-TniQ fusionAttorney Docket No. MTGNZ000100WO comprises a polypeptide having a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion comprises a polypeptide having a sequence identical to SEQ ID NO: 257-307 and 1138-1242. In some embodiments, the FD- TniQ fusion comprises a polypeptide comprising a sequence having at least about 70% identity to SEQ ID NOs: 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion comprises a polypeptide comprising a sequence having at least about 75% identity to SEQ ID NOs: 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion comprises a polypeptide comprising a sequence having at least about 80% identity to SEQ ID NOs: 257- 307 and 1 138-1242. In some embodiments, the FD-TniQ fusion comprises a polypeptide comprising a sequence having at least about 85% identity to SEQ ID NOs: 257-307 and 1138- 1242. In some embodiments, the FD-TniQ fusion comprises a polypeptide comprising a sequence having at least about 90% identity to SEQ ID NOs: 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion comprises a polypeptide comprising a sequence having at least about 91% identity to SEQ ID NOs: 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion comprises a polypeptide comprising a sequence having at least about 92% identity to SEQ ID NOs: 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion comprises a polypeptide comprising a sequence having at least about 93% identity to SEQ ID NOs: 257-307 and 1 138-1242. In some embodiments, the FD-TniQ fusion comprises a polypeptide comprising a sequence having at least about 94% identity to SEQ ID NOs: 257- 307 and 1138-1242. In some embodiments, the FD-TniQ fusion comprises a polypeptide comprising a sequence having at least about 95% identity to SEQ ID NOs: 257-307 and 1138- 1242. In some embodiments, the FD-TniQ fusion comprises a polypeptide comprising a sequence having at least about 96% identity to SEQ ID NOs: 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion comprises a polypeptide comprising a sequence having at least about 97% identity to SEQ ID NOs: 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion comprises a polypeptide comprising a sequence having at least about 98% identity to SEQ ID NOs: 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion comprises a polypeptide comprising a sequence having at least about 99% identity to SEQ ID NOs: 257-307 and 1138-1242. In some embodiments, the FD-TniQ fusion comprises aAttorney Docket No. MTGNZ000100WO polypeptide comprising a sequence having 100% identity to SEQ ID NOs: 257-307 and 1138- 1242.

[0134] In some embodiments, the Tn7 type transposase complex comprises at least one polypeptide comprising a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises a polypeptide comprising a sequence having at least about 70% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises a polypeptide comprising a sequence having at least about 75% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-1 1 1. In some embodiments, the Tn7 type transposase complex comprises a polypeptide comprising a sequence having at least about 80% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises a polypeptide comprising a sequence having at least about 85% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises a polypeptide comprising a sequence having at least about 90% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises a polypeptide comprising a sequence having at least about 91% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-1 11 . In some embodiments, the Tn7 type transposase complex comprises a polypeptide comprising a sequence having at least about 92% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises a polypeptide comprising a sequence having at least about 93% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises a polypeptide comprising a sequence having at least about 94% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises a polypeptide comprising a sequence having at least about 95% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises a polypeptide comprising a sequence having at least about 96% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises a polypeptide comprising a sequence having at least about 97% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, andAttorney Docket No. MTGNZ000100WO109-11 1. In some embodiments, the Tn7 type transposase complex comprises a polypeptide comprising a sequence having at least about 98% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises a polypeptide comprising a sequence having at least about 99% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-11 1. In some embodiments, the Tn7 type transposase complex comprises a polypeptide comprising a sequence having 100% identity to SEQ ID NOs: 3-4, 14- 15, 18-19, 66-67, and 109-111.

[0135] In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide each independently comprising a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111 In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide each independently comprising a sequence with at least 70% sequence identity to any one of SEQ ID NOs: 3-4, 14- 15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide each independently comprising a sequence with at least 75% sequence identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide each independently comprising a sequence with at least 80% sequence identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide each independently having at least about 85% identity to SEQ ID NOs: 3- 4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide each independently having at least about 90% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide each independently having at least about 91% identity to SEQ ID NOs: 3- 4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide each independently having at least about 92% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and aAttorney Docket No. MTGNZ000100WO second polypeptide each independently having at least about 93% identity to SEQ ID NOs: 3- 4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide each independently having at least about 94% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide each independently having at least about 95% identity to SEQ ID NOs: 3- 4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide each independently having at least about 96% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide each independently having at least about 97% identity to SEQ ID NOs: 3- 4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide each independently having at least about 98% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-1 1 1 . In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide each independently having at least about 99% identity to SEQ ID NOs: 3- 4, 14-15, 18-19, 66-67, and 109-111. In some embodiments, the Tn7 type transposase complex comprises at least a first polypeptide and a second polypeptide each independently having 100% identity to SEQ ID NOs: 3-4, 14-15, 18-19, 66-67, and 109-111.

[0136] In some embodiments, the Tn7 type transposase complex comprises an accessory protein. In some embodiments, the accessory protein is ClpX.

[0137] In some embodiments, the accessory protein comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 228-230 and 235-249. In some embodiments, the accessory protein comprises at least one polypeptide comprising a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 228-230 and 235-249. In some embodiments, the accessory protein comprises a polypeptide comprising a sequence having at least about 70% identity to SEQ ID NOs: 228-230 and 235-249. In some embodiments, the accessory protein comprises a polypeptide comprising a sequence having at least about 75% identity to SEQ ID NOs: 228-230 and 235-249. In some embodiments, the accessory protein comprises a polypeptide comprising a sequence having at least about 80% identity to SEQ ID NOs: 228-Attorney Docket No. MTGNZ000100WO230 and 235-249. In some embodiments, the accessory protein comprises a polypeptide comprising a sequence having at least about 85% identity to SEQ ID NOs: 228-230 and 235- 249. In some embodiments, the accessory protein comprises a polypeptide comprising a sequence having at least about 90% identity to SEQ ID NOs: 228-230 and 235-249. In some embodiments, the accessory protein comprises a polypeptide comprising a sequence having at least about 91% identity to SEQ ID NOs: 228-230 and 235-249. In some embodiments, the accessory protein comprises a polypeptide comprising a sequence having at least about 92% identity to SEQ ID NOs: 228-230 and 235-249. In some embodiments, the accessory protein comprises a polypeptide comprising a sequence having at least about 93% identity to SEQ ID NOs: 228-230 and 235-249. In some embodiments, the accessory protein comprises a polypeptide comprising a sequence having at least about 94% identity to SEQ ID NOs: 228- 230 and 235-249. In some embodiments, the accessory protein comprises a polypeptide comprising a sequence having at least about 95% identity to SEQ ID NOs: 228-230 and 235- 249. In some embodiments, the accessory protein comprises a polypeptide comprising a sequence having at least about 96% identity to SEQ ID NOs: 228-230 and 235-249. In some embodiments, the accessory protein comprises a polypeptide comprising a sequence having at least about 97% identity to SEQ ID NOs: 228-230 and 235-249. In some embodiments, the accessory protein comprises a polypeptide comprising a sequence having at least about 98% identity to SEQ ID NOs: 228-230 and 235-249. In some embodiments, the accessory protein comprises a polypeptide comprising a sequence having at least about 99% identity to SEQ ID NOs: 228-230 and 235-249. In some embodiments, the accessory protein comprises a polypeptide comprising a sequence having 100% identity to SEQ ID NOs: 228-230 and 235- 249.

[0138] In some embodiments, the cargo nucleotide sequence is flanked by a left-hand transposase recognition sequence. In some embodiments, the cargo nucleotide sequence is flanked by a right-hand transposase recognition sequence. In some embodiments, the cargo nucleotide sequence is flanked by a left-hand transposase recognition sequence and a righthand transposase recognition sequence.

[0139] In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In someAttorney Docket No. MTGNZ000100WO embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 70% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 75% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 80% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 85% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 90% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 91% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 92% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 93% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 94% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 95% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 96% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 97% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 98% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 99% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having 100% identity to SEQ ID NO: 1404.

[0140] In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 70% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposaseAttorney Docket No. MTGNZ000100WO recognition sequence comprises a nucleotide sequence having at least about 75% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 80% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 85% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 90% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 91% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 92% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 93% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 94% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 95% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 96% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 97% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 98% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 99% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having 100% identity to SEQ ID NO: 1405.

[0141] In some embodiments, the left-hand transposase recognition sequence comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some embodiments, the left-hand transposase recognition sequence comprises a sequence having at least about 70% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some embodiments, the left-hand transposase recognition sequence comprises a sequence having at least about 75% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In someAttorney Docket No. MTGNZ000100WO embodiments, the left-hand transposase recognition sequence comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some embodiments, the left-hand transposase recognition sequence comprises a sequence having at least about 85% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some embodiments, the left-hand transposase recognition sequence comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some embodiments, the left-hand transposase recognition sequence comprises a sequence having at least about 91% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some embodiments, the left-hand transposase recognition sequence comprises a sequence having at least about 92% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some embodiments, the left-hand transposase recognition sequence comprises a sequence having at least about 93% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some embodiments, the left-hand transposase recognition sequence comprises a sequence having at least about 94% identity to any one of SEQ ID NOs: 9, 11 , 36-38, 76, and 78. In some embodiments, the left-hand transposase recognition sequence comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some embodiments, the left-hand transposase recognition sequence comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some embodiments, the left-hand transposase recognition sequence comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some embodiments, the left-hand transposase recognition sequence comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some embodiments, the left-hand transposase recognition sequence comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some embodiments, the left-hand transposase recognition sequence comprises a sequence having 100% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78.

[0142] In some embodiments, the right-hand transposase recognition sequence comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-hand transposase recognition sequence comprises a sequence having at least about 70% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In someAttorney Docket No. MTGNZ000100WO embodiments, the right-hand transposase recognition sequence comprises a sequence having at least about 75% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-hand transposase recognition sequence comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-hand transposase recognition sequence comprises a sequence having at least about 85% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-hand transposase recognition sequence comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-hand transposase recognition sequence comprises a sequence having at least about 91% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-hand transposase recognition sequence comprises a sequence having at least about 92% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-hand transposase recognition sequence comprises a sequence having at least about 93% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-hand transposase recognition sequence comprises a sequence having at least about 94% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-hand transposase recognition sequence comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-hand transposase recognition sequence comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-hand transposase recognition sequence comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-hand transposase recognition sequence comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-hand transposase recognition sequence comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some embodiments, the right-hand transposase recognition sequence comprises a sequence having 100% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93.

[0143] In some embodiments, a system disclosed herein comprises at least one engineered guide polynucleotide, e.g., a gRNA.

[0144] In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least aboutAttorney Docket No. MTGNZ000100WO90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46- 80 consecutive nucleotides identical to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 70% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 75% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 80% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 85% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 90% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 91% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 92% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 93% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 94% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 95% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 96% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80Attorney Docket No. MTGNZ000100WO consecutive nucleotides having at least about 97% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 98% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 99% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having 100% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411.

[0145] In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91 %, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence identical to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 70% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 75% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 80% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 85% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 90% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 91% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 92% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 93% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 94% identity toAttorney Docket No. MTGNZ000100WOSEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 95% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 96% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 97 % identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 98% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 99% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411. In some embodiments the engineered guide polynucleotide is a guide RNA comprises a sequence having 100% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411.

[0146] In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754- 944. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides identical to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 70% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 75% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754- 944. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having atAttorney Docket No. MTGNZOOOIOOWO least about 85% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119- 140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 90% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94- 108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 91% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 92% identity to any one of SEQ ID NOs: 5-6, 32- 33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46- 80 consecutive nucleotides having at least about 93% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94- 108, 119- 140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 94% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 95% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 96% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 97% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 98% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 99% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 consecutive nucleotides having 100% identity to any one of SEQ ID NOs: 5-6, 32- 33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944.Attorney Docket No. MTGNZ000100WO

[0147] In some embodiments, the engineered guide polynucleotide is a guide RNA and comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence identical to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 70% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 75% identity to any one of SEQ ID NOs: 5-6, 32- 33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 85% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94- 108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 91% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 92% identity to any one of SEQ ID NOs: 5-6, 32- 33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 93% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 94% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94- 108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least aboutAttorney Docket No. MTGNZ000100WO96% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 5-6, 32- 33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944. In some embodiments, the engineered guide polynucleotide is a guide RNA comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 5-6, 32-33, 45-63, 68-75, 94- 108, 119-140, 222, and 754-944. In some embodiments the engineered guide polynucleotide is a guide RNA comprises a sequence having 100% identity to any one of SEQ ID NOs: 5-6, 32- 33, 45-63, 68-75, 94-108, 119-140, 222, and 754-944.

[0148] In some embodiments, the guide RNAs comprise various structural elements including but not limited to: a spacer sequence which binds to the protospacer sequence (target sequence), a crRNA, and an optional tracrRNA. In some embodiments, the guide RNA comprises a crRNA comprising a spacer sequence. In some embodiments, the guide RNA additionally comprises a tracrRNA or a modified tracrRNA.

[0149] In some embodiments, the systems provided herein comprise one or more guide RNAs. In some embodiments, the guide RNA comprises a sense sequence. In some embodiments, the guide RNA comprises an anti-sense sequence. In some embodiments, the guide RNA comprises nucleotide sequences other than the region complementary to or substantially complementary to a region of a target sequence. For example, a crRNA is part or considered part of a guide RNA, or is comprised in a guide RNA, e.g., a crRNA: tracrRNA chimera.

[0150] In some embodiments, the guide RNA comprises synthetic nucleotides or modified nucleotides. In some embodiments, the guide RNA comprises one or more inter-nucleoside linkers modified from the natural phosphodiester. In some embodiments, all of the inter- nucleoside linkers of the guide RNA, or contiguous nucleotide sequence thereof, are modified. For example, in some embodiments, the inter nucleoside linkage comprises Sulphur (S), such as a phosphorothioate inter-nucleoside linkage.

[0151] In some embodiments, the guide RNA comprises modifications to a ribose sugar or nucleobase. In some embodiments, the guide RNA comprises one or more nucleosides comprising a modified sugar moiety, wherein the modified sugar moiety is a modification of the sugar moiety when compared to the ribose sugar moiety found in deoxyribose nucleic acid (DNA) and RNA. In some embodiments, the modification is within the ribose ring structure. Exemplary modifications include, but are not limited to, replacement with a hexose ring(HNA), a bicyclic ring having a biradical bridge between the C2 and C4 carbons on the riboseAttorney Docket No. MTGNZ000100WO ring (e.g., locked nucleic acids (LNA)), or an unlinked ribose ring which typically lacks a bond between the C2 and C3 carbons (e.g., UNA). In some embodiments, the sugar-modified nucleosides comprise bicyclohexose nucleic acids or tricyclic nucleic acids. In some embodiments, the modified nucleosides comprise nucleosides where the sugar moiety is replaced with a non-sugar moiety, for example peptide nucleic acids (PNA) or morpholino nucleic acids.

[0152] In some embodiments, the guide RNA comprises one or more modified sugars. In some embodiments, the sugar modifications comprise modifications made by altering the substituent groups on the ribose ring to groups other than hydrogen, or the 2 ’-OH group naturally found in DNA and RNA nucleosides. In some embodiments, substituents are introduced at the 2’, 3’, 4’, or 5’ positions, or combinations thereof. In some embodiments, nucleosides with modified sugar moieties comprise 2’ modified nucleosides, e.g., 2’ substituted nucleosides. A 2’ sugar modified nucleoside, in some embodiments, is a nucleoside that has a substituent other than -H or -OH at the 2’ position (2’ substituted nucleoside) or comprises a 2’ linked biradical, and comprises 2’ substituted nucleosides and LNA (2’-4’ biradical bridged) nucleosides. Examples of 2’ -substituted modified nucleosides comprise, but are not limited to, 2’-O-alkyl-RNA, 2’-O- methyl-RNA, 2’-alkoxy-RNA, 2’ -O-methoxyethyl-RNA (MOE), 2’-amino-DNA, 2’-Fluoro- RNA, and 2’-F-ANA nucleosides. In some embodiments, the modification in the ribose group comprises a modification at the 2’ position of the ribose group. In some embodiments, the modification at the 2’ position of the ribose group is selected from the group consisting of 2’- O-methyl, 2’-fluoro, 2’-deoxy, and 2’-O-(2-methoxyethyl).

[0153] In some embodiments, the guide RNA comprises one or more modified sugars. In some embodiments, the guide RNA comprises only modified sugars. In certain embodiments, the guide RNA comprises greater than about 10%, 25%, 50%, 75%, or 90% modified sugars. In some embodiments, the modified sugar is a bicyclic sugar. In some embodiments, the modified sugar comprises a 2’-O-methoxyethyl group. In some embodiments, the guide RNA comprises both inter-nucleoside linker modifications and nucleoside modifications.

[0154] In some embodiments, the guide RNA comprises a sequence complementary to a eukaryotic, fungal, plant, mammalian, or human genomic polynucleotide sequence. In some embodiments, the guide RNA comprises a sequence complementary to a eukaryotic genomic polynucleotide sequence. In some embodiments, the guide RNA comprises a sequence complementary to a fungal genomic polynucleotide sequence. In some embodiments, the guide RNA comprises a sequence complementary to a plant genomic polynucleotide sequence. In some embodiments, the guide RNA comprises a sequence complementary to a mammalianAttorney Docket No. MTGNZ000100WO genomic polynucleotide sequence. In some embodiments, the guide RNA comprises a sequence complementary to a human genomic polynucleotide sequence.

[0155] In some embodiments, the guide RNA is 30-250 nucleotides in length. In some embodiments, the guide RNA is more than 90 nucleotides in length. In some embodiments, the guide RNA is less than 245 nucleotides in length. In some embodiments, the guide RNA is 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 220, 240, or more than 240 nucleotides in length. In some embodiments, the guide RNA is about 30 to about 40, about 30 to about 50, about 30 to about 60, about 30 to about 70, about 30 to about 80, about 30 to about 90, about 30 to about 100, about 30 to about 120, about 30 to about 140, about 30 to about 160, about 30 to about 180, about 30 to about 200, about 30 to about 220, about 30 to about 240, about 50 to about 60, about 50 to about 70, about 50 to about 80, about 50 to about 90, about 50 to about 100, about 50 to about 120, about 50 to about 140, about 50 to about 160, about 50 to about 180, about 50 to about 200, about 50 to about 220, about 50 to about 240, about 100 to about 120, about 100 to about 140, about 100 to about 160, about 100 to about 180, about 100 to about 200, about 100 to about 220, about 100 to about 240, about 160 to about 180, about 160 to about 200, about 160 to about 220, or about 160 to about 240 nucleotides in length.

[0156] In some embodiments, the system further comprises a PAM sequence compatible with the Cas effector complex. In some embodiments, the PAM sequence comprises SEQ ID NO: 31.

[0157] In some embodiments, the PAM sequence is located about 50 to about 70 base pairs from the target nucleic acid site. In some embodiments, the PAM sequence is located 3’ of the target nucleic acid site. In some embodiments, the PAM sequence is located 5’ of the target nucleic acid site.

[0158] In some embodiments, the class 2, type V effector comprises a nuclear localization sequence (NLS). In some embodiments, the NLS is at an N-terminus of the class 2, type V effector. In some embodiments, the NLS is at a C-terminus of the class 2, type V effector. In some embodiments, the NLS is at an N-terminus and a C-terminus of the class 2, type V effector.

[0159] In some embodiments, the NLS comprises a sequence of any one of SEQ ID NOs: 192- 207 and 1354-1383, or a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 192-207Attorney Docket No. MTGNZ000100WO and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 80% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 85% identity to SEQ ID NOs: 192-207 and 1354- 1383. In some embodiments, the NLS comprises a sequence having at least about 90% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 91% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 92% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 93% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 94% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 95% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 96% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 97% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 98% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 99% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having 100% identity to SEQ ID NOs: 192-207 and 1354-1383.Table 1: Exemplary NLS SequencesAttorney Docket No. MTGNZOOOIOOWOAttorney Docket No. MTGNZ000100WO

[0160] In some embodiments, the Cas effector complex further comprises a small prokaryotic ribosomal protein subunit SI 5. In some embodiments, the SI 5 fusion protein is encoded by a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 181-183. In some embodiments, the S 15 is encoded by a sequence having at least about 70% identity to SEQ ID NOs: 181 -183. In some embodiments, the S15 is encoded by a sequence having at least about 75% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having at least about 80% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having at least about 85% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having at least about 90% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having at least about 91% identity to SEQ ID NOs: 181-183. In some embodiments, the S 15 is encoded by a sequence having at least about 92% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having at least about 93% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having at least about 94% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having at least about 95% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having at least about 96% identity to SEQ ID NOs: 181-183. In someAttorney Docket No. MTGNZ000100WO embodiments, the S15 is encoded by a sequence having at least about 97% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having at least about 98% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having at least about 99% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having 100% identity to SEQ ID NOs: 181-183.

[0161] In some embodiments, the S15 comprises a sequence having at least about 70% identity to SEQ ID NOs: 187-189. In some embodiments, the S15 comprises a sequence having at least about 75% identity to SEQ ID NOs: 187-189. In some embodiments, the S 15 comprises a sequence having at least about 80% identity to SEQ ID NOs: 187-189. In some embodiments, the S15 comprises a sequence having at least about 85% identity to SEQ ID NOs: 187-189. In some embodiments, the S15 comprises a sequence having at least about 90% identity to SEQ ID NOs: 187-189. In some embodiments, the S15 comprises a sequence having at least about 91% identity to SEQ ID NOs: 187-189. In some embodiments, the S 15 comprises a sequence having at least about 92% identity to SEQ ID NOs: 187-189. In some embodiments, the SI 5 comprises a sequence having at least about 93% identity to SEQ ID NOs: 187-189. In some embodiments, the S15 comprises a sequence having at least about 94% identity to SEQ ID NOs: 187-189. In some embodiments, the S15 comprises a sequence having at least about 95% identity to SEQ ID NOs: 187-189. In some embodiments, the S 15 comprises a sequence having at least about 96% identity to SEQ ID NOs: 187-189. In some embodiments, the S15 comprises a sequence having at least about 97% identity to SEQ ID NOs: 187-189. In some embodiments, the S15 comprises a sequence having at least about 98% identity to SEQ ID NOs: 187-189. In some embodiments, the S 15 comprises a sequence having at least about 99% identity to SEQ ID NOs: 187-189. In some embodiments, the S 15 comprises a sequence having 100% identity to SEQ ID NOs: 187-189.

[0162] In some embodiments, the Cas effector complex comprises one or more linkers linking the class 2, type V effector, the small prokaryotic ribosomal protein subunit S 15, the transposase, the gRNA, or combinations thereof. In some embodiments, the linker comprises at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, or 400 amino acids. In some embodiments, the linker comprises at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides. In some embodiments, the linker is encoded by a sequence of SEQ ID NO: 186, or a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at leastAttorney Docket No. MTGNZ000100WO about 96%, at least about 97%, at least about 98%, or at least about 99% identity of SEQ ID NO: 186. In some embodiments, the linker is encoded by SEQ ID NO: 186.

[0163] In some aspects, the present disclosure provides an engineered nuclease system comprising: an endonuclease comprising a RuvC domain, the endonuclease being derived from an uncultivated microorganism and is a Class 2, type V-K Cas effector comprising at least 80% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 220; and an engineered guide RNA that forms a complex with the endonuclease and comprising a spacer sequence that hybridizes to a target nucleic acid sequence wherein the engineered guide polynucleotide comprises a sequence comprising at least 80% identity to any one of SEQ ID NOs: 754-944.Fusion Proteins

[0164] Described herein, in some embodiments, are systems for transposing a cargo nucleotide sequence into a target nucleic acid site comprising a fusion protein or a nucleic acid encoding the fusion protein. In some embodiments, the fusion protein or a nucleic acid encoding the fusion protein comprises a class 2, type V effector, a small prokaryotic ribosomal protein subunit S15, a transposase, a gRNA, or combinations thereof. In some embodiments, the fusion protein comprises one or more transposases.

[0165] In some embodiments, a nuclear localization sequence (NLS) is fused to the class 2, type V effector. In some embodiments, the NLS is fused at an N-terminus of the class 2, type V effector. In some embodiments, the NLS is fused at a C-terminus of the class 2, type V effector. In some embodiments, the NLS is fused at an N-terminus and a C-terminus of the class 2, type V effector.

[0166] In some embodiments, the NLS comprises a sequence of any one of SEQ ID NOs: 192- 207 and 1354-1383, or a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 80% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 85% identity to SEQ ID NOs: 192-207 and 1354- 1383. In some embodiments, the NLS comprises a sequence having at least about 90% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 91% identity to SEQ ID NOs: 192-207 and 1354-1383. In someAttorney Docket No. MTGNZ000100WO embodiments, the NLS comprises a sequence having at least about 92% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 93% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 94% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 95% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 96% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 97% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 98% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having at least about 99% identity to SEQ ID NOs: 192-207 and 1354-1383. In some embodiments, the NLS comprises a sequence having 100% identity to SEQ ID NOs: 192-207 and 1354-1383.

[0167] In some embodiments, the fusion protein or a nucleic acid encoding the fusion protein comprises a fusion of S15 and a nuclear localization sequence (NLS). In some embodiments, the NLS is fused at an N-terminus of SI 5. In some embodiments, the NLS is fused at a C- terminus of SI 5. In some embodiments, the NLS is fused at an N-terminus and a C-terminus of S15.

[0168] In some embodiments, the S15 fusion protein further comprises a cleavable peptide. In some embodiments, the peptide is a 2 A peptide.

[0169] In some embodiments, the S 15 fusion protein is encoded by a sequence with at least 80% sequence identity to any one of SEQ ID NOs: 181-183. In some embodiments, the S15 fusion protein is encoded by a sequence with at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 181- 183. In some embodiments, the Cas effector complex further comprises a small prokaryotic ribosomal protein subunit SI 5. In some embodiments, the S15 fusion protein is encoded by a sequence with at least 80% sequence identity to any one of SEQ ID NOs: 181-183. In some embodiments, the S15 fusion protein is encoded by a sequence with at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least aboutAttorney Docket No. MTGNZ000100WO91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having at least about 70% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having at least about 75% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having at least about 80% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having at least about 85% identity to SEQ ID NOs: 181-183. In some embodiments, the S 15 is encoded by a sequence having at least about 90% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having at least about 91% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having at least about 92% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having at least about 93% identity to SEQ ID NOs: 181-183. In some embodiments, the S 15 is encoded by a sequence having at least about 94% identity to SEQ ID NOs: 181- 183. In some embodiments, the S15 is encoded by a sequence having at least about 95% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having at least about 96% identity to SEQ ID NOs: 181-183. In some embodiments, the S 15 is encoded by a sequence having at least about 97% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having at least about 98% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having at least about 99% identity to SEQ ID NOs: 181-183. In some embodiments, the S15 is encoded by a sequence having 100% identity to SEQ ID NOs: 181-183.

[0170] In some embodiments, the S15 fusion protein comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 187-189. In some embodiments, the S15 fusion protein has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 187-189. In some embodiments, the S15 comprises a sequence having at least about 70% identity to SEQ ID NOs: 187-189. In some embodiments, the S15 comprises a sequence having at least about 75% identity to SEQ ID NOs: 187-189. In some embodiments, the S15 comprises a sequence having at least about 80% identity to SEQ ID NOs: 187-189. In some embodiments, the S15 comprises a sequence having at least about 85% identity to SEQ ID NOs: 187-189. In someAttorney Docket No. MTGNZ000100WO embodiments, the S 15 comprises a sequence having at least about 90% identity to SEQ ID NOs: 187-189. In some embodiments, the S 15 comprises a sequence having at least about 91% identity to SEQ ID NOs: 187-189. In some embodiments, the S15 comprises a sequence having at least about 92% identity to SEQ ID NOs: 187-189. In some embodiments, the S15 comprises a sequence having at least about 93% identity to SEQ ID NOs: 187-189. In some embodiments, the S15 comprises a sequence having at least about 94% identity to SEQ ID NOs: 187-189. In some embodiments, the S15 comprises a sequence having at least about 95% identity to SEQ ID NOs: 187-189. In some embodiments, the S15 comprises a sequence having at least about 96% identity to SEQ ID NOs: 187-189. In some embodiments, the S15 comprises a sequence having at least about 97% identity to SEQ ID NOs: 187-189. In some embodiments, the S15 comprises a sequence having at least about 98% identity to SEQ ID NOs: 187-189. In some embodiments, the S15 comprises a sequence having at least about 99% identity to SEQ ID NOs: 187-189. In some embodiments, the S 15 comprises a sequence having 100% identity to SEQ ID NOs: 187-189.

[0171] In some embodiments, an NLS is fused to the transposase. In some embodiments, the transposase is TnsB, TnsC, or TniQ. In some embodiments, the transposase is TnsB. In some embodiments, the transposase is TnsC. In some embodiments, the transposase is TniQ. In some embodiments, the NLS is fused at an N-terminus of the transposase. In some embodiments, the NLS is fused at a C-terminus of the transposase. In some embodiments, the NLS is fused at an N-terminus and a C-terminus of the transposase.

[0172] In some embodiments, the fusion protein or a nucleic acid encoding the fusion protein comprises a gRNA described herein (for example a dual gRNA or a single gRNA).

[0173] In some embodiments, the class 2, type V effector, the small prokaryotic ribosomal protein subunit S15, the transposase, the gRNA, or a fusion protein comprises a tag. In some embodiments, the tag is an affinity tag. In some embodiments the tag is a polypeptide or a polynucleotide. Exemplary affinity tags include, but are not limited to, a His-tag, a Flag tag, a Myc-tag, an MBP-tag, and a GST-tag.

[0174] In some embodiments, the class 2, type V effector, the small prokaryotic ribosomal protein subunit SI 5, the transposase, or a fusion protein, comprises a protease cleavage site. Exemplary protease cleavage sites include, but are not limited to, a TEV site, a C3 site, a Factor Xa site, and an Enterokinase site.Cells

[0175] Described herein, in certain embodiments, is a cell comprising the systems described herein.Attorney Docket No. MTGNZOOOIOOWO

[0176] In some embodiments, the cell is a eukaryotic cell (e.g., a plant cell, an animal cell, a protist cell, or a fungi cell), a mammalian cell (a Chinese hamster ovary (CHO) cell, baby hamster kidney (BHK), human embryo kidney (HEK), mouse myeloma (NSO), or human retinal cells), an immortalized cell (e.g., a HeLa cell, a COS cell, a HEK-293T cell, a MDCK cell, a 3T3 cell, a PC 12 cell, a Huh7 cell, a HepG2 cell, a K562 cell, a N2a cell, or a SY5Y cell), an insect cell (e.g., a Spodoptera frugiperda cell, a Trichoplusia ni cell, a Drosophila melanogaster cell, a S2 cell, or a Heliothis virescens cell), a yeast cell (e.g., a Saccharomyces cerevisiae cell, a Cryptococcus cell, or a Candida cell), a plant cell (e.g., a parenchyma cell, a collenchyma cell, or a sclerenchyma cell), a fungal cell (e.g., a Saccharomyces cerevisiae cell, a Cryptococcus cell, or a Candida cell), or a prokaryotic cell (e.g., a E. coli cell, a streptococcus bacterium cell, a streptomyces soil bacteria cell, or an archaea cell). In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is an immortalized cell. In some embodiments, the cell is an insect cell. In some embodiments, the cell is a yeast cell. In some embodiments, the cell is a plant cell. In some embodiments, the cell is a fungal cell. In some embodiments, the cell is a prokaryotic cell.

[0177] In some embodiments, the cell is an A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1 , Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cell, HT1080, HepG2, Huh7, K562, a primary cell, or derivative thereof.Delivery and Vectors

[0178] Disclosed herein, in some embodiments, are nucleic acid sequences encoding a MG64 system comprising a class 2, type V effector, a small prokaryotic ribosomal protein subunit SI 5, a transposase, a gRNA, a fusion protein or a gene editing system disclosed herein.

[0179] In some embodiments, the nucleic acid encoding the MG64 system is a DNA, for example a linear DNA, a plasmid DNA, or a minicircle DNA. In some embodiments, the nucleic acid encoding the MG64 system is an RNA, for example a mRNA.

[0180] In some embodiments, the nucleic acid encoding the MG64 system is delivered by a nucleic acid-based vector. In some embodiments, the nucleic acid-based vector is a plasmid (e.g., circular DNA molecules that can autonomously replicate inside a cell), cosmid (e.g., pWE or sCos vectors), artificial chromosome, human artificial chromosome (HAC), yeast artificial chromosomes (YAC), bacterial artificial chromosome (BAC), Pl -derived artificial chromosomes (PAC), phagemid, phage derivative, bacmid, or virus. In some embodiments, the nucleic acid-based vector is selected from the list consisting of: pSF-CMV-NEO-NH2-PPT- 3XFLAG, pSF-CMV-NEO-COOH-3XFLAG, pSF-CMV-PURO-NH2-GST-TEV, pSF-Attorney Docket No. MTGNZ000100WOOXB20-COOH-TEV-FLAG(R)-6His, pCEP4 pDEST27, pSF-CMV-Ub-KrYFP, pSF-CMV- FMDV-daGFP, pEFla-mCherry-Nl vector, pEFla-tdTomato vector, pSF-CMV-FMDV- Hygro, pSF-CMV-PGK-Puro, pMCP-tag(m), pSF-CMV-PUR0-NH2-CMYC, pSF-OXB20- BetaGal,pSF-OXB20-Fhic, pSF-OXB20, pSF-Tac, pRI 101 -AN DNA, pCambia2301,pTYB21, pKLAC2, pAc5.1 / V5-His A, and pDEST8.

[0181] In some embodiments, the nucleic acid-based vector comprises a promoter. In some embodiments, the promoter is selected from the group consisting of a mini promoter, an inducible promoter, a constitutive promoter, and derivatives thereof. In some embodiments, the promoter is selected from the group consisting of CMV, CBA, EFla, CAG, PGK, TRE, U6, UAS, T7, Sp6, lac, araBad, trp, Ptac, p5, pl9, p40, Synapsin, CaMKII, GRK1, and derivatives thereof. In some embodiments the promoter is a U6 promoter. In some embodiments, the promoter is a CAG promoter. In some embodiments, the promoter is encoded by a sequence of any one of SEQ ID NOs: 190-191, or a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity of any one of SEQ ID NOs: 190-191.

[0182] In some embodiments, the nucleic acid-based vector is a virus. In some embodiments, the virus is an alphavirus, a parvovirus, an adenovirus, an AAV, a baculovirus, a Dengue virus, a lentivirus, a herpesvirus, a poxvirus, an anellovirus, a bocavirus, a vaccinia virus, or a retrovirus. In some embodiments, the virus is an alphavirus. In some embodiments, the virus is a parvovirus. In some embodiments, the virus is an adenovirus. In some embodiments, the virus is an AAV. In some embodiments, the virus is a baculovirus. In some embodiments, the virus is a Dengue virus. In some embodiments, the virus is a lentivirus. In some embodiments, the virus is a herpesvirus. In some embodiments, the virus is a poxvirus. In some embodiments, the virus is an anellovirus. In some embodiments, the virus is a bocavirus. In some embodiments, the virus is a vaccinia virus. In some embodiments, the virus is or a retrovirus.

[0183] In some embodiments, the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV 11, AAV12, AAV13, AAV14, AAV 15, AAV16, AAV- rh8, AAV-rhlO, AAV-rh20, AAV-rh39, AAV-rh74, AAV-rhM4-l, AAV-hu37, AAV-Anc80, AAV-Anc80L65, AAV-7m8, AAV-PHP-B, AAV-PHP-EB, AAV-2.5, AAV-2tYF, AAV-3B, AAV-LK03, AAV-HSC1, AAV-HSC2, AAV-HSC3, AAV-HSC4, AAV-HSC5, AAV-HSC6, AAV-HSC7, AAV-HSC8, AAV-HSC9, AAV-HSC10, AAV-HSC11, AAV-HSC12, AAV-Attorney Docket No. MTGNZOOOIOOWOHSC13, AAV-HSC14, AAV-HSC15, AAV-TT, AAV-DJ / 8, AAV-Myo, AAV-NP40, AAV- NP59, AAV-NP22, AAV-NP66, AAV-HSC16, or a derivative thereof. In some embodiments, the herpesvirus is HSV type 1, HSV-2, VZV, EBV, CMV, HHV-6, HHV-7, or HHV-8.

[0184] In some embodiments, the virus is AAV 1 or a derivative thereof. In some embodiments, the virus is AAV2 or a derivative thereof. In some embodiments, the virus is AAV3 or a derivative thereof. In some embodiments, the virus is AAV4 or a derivative thereof. In some embodiments, the virus is AAV5 or a derivative thereof. In some embodiments, the virus is AAV6 or a derivative thereof. In some embodiments, the virus is AAV7 or a derivative thereof. In some embodiments, the virus is AAV8 or a derivative thereof. In some embodiments, the virus is AAV9 or a derivative thereof. In some embodiments, the virus is AAV10 or a derivative thereof. In some embodiments, the virus is AAV 11 or a derivative thereof. In some embodiments, the virus is AAV 12 or a derivative thereof. In some embodiments, the virus is AAV13 or a derivative thereof. In some embodiments, the virus is A AVI 4 or a derivative thereof. In some embodiments, the virus is AAV 15 or a derivative thereof. In some embodiments, the virus is AAV 16 or a derivative thereof. In some embodiments, the virus is AAV-rh8 or a derivative thereof. In some embodiments, the virus is AAV-rhlO or a derivative thereof. In some embodiments, the virus is AAV-rh20 or a derivative thereof. In some embodiments, the virus is AAV-rh39 or a derivative thereof. In some embodiments, the virus is AAV-rh74 or a derivative thereof. In some embodiments, the virus is AAV-rhM4-l or a derivative thereof. In some embodiments, the virus is AAV-hu37 or a derivative thereof. In some embodiments, the virus is AAV-Anc80 or a derivative thereof. In some embodiments, the virus is AAV-Anc80L65 or a derivative thereof. In some embodiments, the virus is AAV-7m8 or a derivative thereof. In some embodiments, the virus is AAV-PHP-B or a derivative thereof. In some embodiments, the virus is AAV-PHP-EB or a derivative thereof. In some embodiments, the virus is AAV-2.5 or a derivative thereof. In some embodiments, the virus is AAV-2tYF or a derivative thereof. In some embodiments, the virus is AAV-3B or a derivative thereof. In some embodiments, the virus is AAV-LK03 or a derivative thereof. In some embodiments, the virus is AAV-HSC1 or a derivative thereof. In some embodiments, the virus is AAV-HSC2 or a derivative thereof. In some embodiments, the virus is AAV-HSC3 or a derivative thereof. In some embodiments, the virus is AAV-HSC4 or a derivative thereof. In some embodiments, the virus is AAV-HSC5 or a derivative thereof. In some embodiments, the virus is AAV-HSC6 or a derivative thereof. In some embodiments, the virus is AAV-HSC7 or a derivative thereof. In some embodiments, the virus is AAV-HSC8 or a derivative thereof. In some embodiments, the virus is AAV-HSC9 or a derivative thereof. In some embodiments, the virus is AAV-HSC10 or a derivative thereof. InAttorney Docket No. MTGNZOOOIOOWO some embodiments, the virus is AAV-HSC11 or a derivative thereof. In some embodiments, the virus is AAV-HSC12 or a derivative thereof. In some embodiments, the virus is AAV- HSC13 or a derivative thereof. In some embodiments, the virus is AAV-HSC14 or a derivative thereof. In some embodiments, the virus is AAV-HSC15 or a derivative thereof. In some embodiments, the virus is AAV-TT or a derivative thereof. In some embodiments, the virus is AAV-DJ / 8 or a derivative thereof. In some embodiments, the virus is AAV-Myo or a derivative thereof. In some embodiments, the virus is AAV-NP40 or a derivative thereof. In some embodiments, the virus is AAV-NP59 or a derivative thereof. In some embodiments, the virus is AAV-NP22 or a derivative thereof. In some embodiments, the virus is AAV-NP66 or a derivative thereof. In some embodiments, the virus is AAV-HSC16 or a derivative thereof.

[0185] In some embodiments, the virus is HSV-1 or a derivative thereof. In some embodiments, the virus is HSV-2 or a derivative thereof. In some embodiments, the virus is VZV or a derivative thereof. In some embodiments, the virus is EBV or a derivative thereof. In some embodiments, the virus is CMV or a derivative thereof. In some embodiments, the virus is HHV-6 or a derivative thereof. In some embodiments, the virus is HHV-7 or a derivative thereof. In some embodiments, the virus is HHV-8 or a derivative thereof.

[0186] In some embodiments, the nucleic acid encoding the MG64 system is delivered by a non-nucleic acid-based delivery system (e.g., a non-viral delivery system). In some embodiments, the non-viral delivery system is a liposome. In some embodiments, the nucleic acid is associated with a lipid. The nucleic acid associated with a lipid, in some embodiments, is encapsulated in the aqueous interior of a liposome, interspersed within the lipid bilayer of a liposome, attached to a liposome via a linking molecule that is associated with both the liposome and the nucleic acid, entrapped in a liposome, complexed with a liposome, dispersed in a solution containing a lipid, mixed with a lipid, combined with a lipid, contained as a suspension in a lipid, contained or complexed with a micelle, or otherwise associated with a lipid. In some embodiments, the nucleic acid is comprised in a lipid nanoparticle (LNP).

[0187] In some embodiments, the fusion protein or genome editing system is introduced into the cell in any suitable way, either stably or transiently. In some embodiments, a fusion protein or genome editing system is transfected into the cell. In some embodiments, the cell is transduced or transfected with a nucleic acid construct that encodes a fusion protein or genome editing system. For example, a cell is transduced (e.g., with a virus encoding a fusion protein or genome editing system), or transfected (e.g., with a plasmid encoding a fusion protein or genome editing system) with a nucleic acid that encodes a fusion protein or genome editing system, or the translated fusion protein or genome editing system. In some embodiments, the transduction is a stable or transient transduction. In some embodiments, cells expressing aAttorney Docket No. MTGNZ000100WO fusion protein or genome editing system or containing a fusion protein or genome editing system are transduced or transfected with one or more gRNA molecules, for example, when the fusion protein or genome editing system comprises a CRISPR nuclease. In some embodiments, a plasmid expressing a fusion protein or genome editing system is introduced into cells through electroporation, transient (e.g., lipofection) and stable genome integration (e.g., piggybac) and viral transduction (for example lentivirus or AAV) or other methods known to those of skill in the art. In some embodiments, the gene editing system is introduced into the cell as one or more polypeptides. In some embodiments, delivery is achieved through the use of RNP complexes. Delivery methods to cells for polypeptides and / or RNPs are known in the art, for example by electroporation or by cell squeezing.

[0188] Exemplary methods of delivery of nucleic acids include lipofection, nucleofection, electroporation, stable genome integration (e.g., piggybac), microinjection, biolistics, virosomes, liposomes, immunoliposomes, polycation or lipid nucleic acid conjugates, naked DNA, artificial virions, and agent-enhanced uptake of DNA. Lipofection is described in e.g., U.S. Pat. Nos. 5,049,386; 4,946,787; and 4,897,355) and lipofection reagents are sold commercially (e.g., Transfectam™, Lipofectin™ and SF Cell Line 4D-Nucleofector X Kit™ (Lonza)). Cationic and neutral lipids that are suitable for efficient receptor-recognition lipofection of polynucleotides include those of WO 91 / 17424 and WO 91 / 16024. In some embodiments, the delivery is to cells (e.g., in vitro or ex vivo administration) or target tissues (e.g., in vivo administration). In some embodiments, the nucleic acid is comprised in a liposome or a nanoparticle that specifically targets a host cell.

[0189] Additional methods for the delivery of nucleic acids to cells are known to those skilled in the art. See, for example, US 2003 / 0087817.

[0190] In some embodiments, the present disclosure provides a cell comprising a vector or a nucleic acid described herein. In some embodiments, the cell expresses a gene editing system or parts thereof. In some embodiments, the cell is a human cell. In some embodiments, the cell is genome edited ex vivo. In some embodiments, the cell is genome edited in vivo.Methods for Transposition

[0191] The present disclosure provides methods for transposing a cargo nucleotide sequence into a target nucleic acid site. In some embodiments, the method comprises expressing a system described herein within a cell or introducing a system described herein to a cell. In some embodiments, the method comprises contacting a cell with a system described herein.

[0192] In some embodiments, the method comprises contacting a double- stranded nucleic acid comprising the cargo nucleotide sequence with a Cas effector complex comprising a class 2,Attorney Docket No. MTGNZ000100WO type V Cas effector and at least one engineered guide polynucleotide configured to hybridize to the target nucleotide sequence. In some embodiments, the method comprises contacting the double-stranded nucleic acid comprising the cargo nucleotide sequence with a Tn7 type transposase complex configured to bind the Cas effector complex, wherein the Tn7 type transposase complex comprises a TnsB subunit.

[0193] In some embodiments, the method comprises contacting a double- stranded nucleic acid comprising the cargo nucleotide sequence with a Cas effector complex and a Tn7 type transposase complex that binds the Cas effector complex, wherein the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence having at least 70% sequence identity to any one of SEQ ID NOs: 1406-1409.

[0194] In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 70% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 75% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 80% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 85% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 90% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 91% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 92% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 93% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complexAttorney Docket No. MTGNZ000100WO comprise a nucleotide sequence having at least about 94% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 95% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 96% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 97% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 98% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having at least about 99% identity to any one of SEQ ID NOs: 1406-1409. In some embodiments, the Cas effector complex and the Tn7 type transposase complex comprise a nucleotide sequence having 100% identity to any one of SEQ ID NOs: 1406-1409.

[0195] In some embodiments, the cargo nucleotide sequence is flanked by a left-hand transposase recognition sequence. In some embodiments, the cargo nucleotide sequence is flanked by a right-hand transposase recognition sequence. In some embodiments, the cargo nucleotide sequence is flanked by a left-hand transposase recognition sequence and a righthand transposase recognition sequence.

[0196] In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 9, 11, 36-38, 76, and 78. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 70% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 75% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 80% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 85% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 90% identity to SEQ ID NO: 1404. In some embodiments, the left-handAttorney Docket No. MTGNZ000100WO transposase recognition sequence comprises a nucleotide sequence having at least about 91% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 92% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 93% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 94% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 95% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 96% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 97% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 98% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least about 99% identity to SEQ ID NO: 1404. In some embodiments, the left-hand transposase recognition sequence comprises a nucleotide sequence having 100% identity to SEQ ID NO: 1404.

[0197] In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 70% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 75% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 80% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 85% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 90% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 91% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequenceAttorney Docket No. MTGNZOOOIOOWO comprises a nucleotide sequence having at least about 92% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 93% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 94% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 95% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 96% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 97% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 98% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least about 99% identity to SEQ ID NO: 1405. In some embodiments, the right-hand transposase recognition sequence comprises a nucleotide sequence having 100% identity to SEQ ID NO: 1405.

[0198] In some embodiments, the method further comprises a PAM sequence compatible with the Cas effector complex adjacent to the target nucleic acid site. In some embodiments, the PAM sequence is located 3 ’ of the target nucleic acid site.Uses

[0199] Systems of the present disclosure may be used for various applications, such as, for example, nucleic acid editing (e.g., gene editing) or binding to a nucleic acid molecule (e.g., sequence-specific binding). Such systems may be used, for example, for remediating (e.g., removing or replacing) a genetically inherited mutation that may cause a disease in a subject; inactivating a gene in order to ascertain its function in a cell; as a diagnostic tool to detect disease-causing genetic elements (e.g., via cleavage of reverse-transcribed viral RNA or an amplified DNA sequence encoding a disease-causing mutation); as deactivated enzymes in combination with a probe to target and detect a specific nucleotide sequence (e.g., sequence encoding antibiotic resistance int bacteria); to render viruses inactive or incapable of infecting host cells by targeting viral genomes; to add genes or amend metabolic pathways to engineer organisms to produce valuable small molecules, macromolecules, or secondary metabolites; to establish a gene drive element for evolutionary selection, and / or to detect cell perturbations by foreign small molecules and nucleotides as a biosensor.KitsAttorney Docket No. MTGNZOOOIOOWO

[0200] In some embodiments, this disclosure provides kits comprising one or more nucleic acid constructs encoding the various components of the fusion protein or genome editing system described herein, e.g., comprising a nucleotide sequence encoding the components of the fusion protein or genome editing system capable of modifying a target DNA sequence. In some embodiments, the nucleotide sequence comprises a heterologous promoter that drives expression of the RNA genome editing system components.

[0201] In some embodiments, the class 2, type V effector, the small prokaryotic ribosomal protein subunit SI 5, the transposase, the gRNA, or a fusion protein or gene editing system comprising any combination thereof disclosed herein is assembled into a pharmaceutical, diagnostic, or research kit to facilitate its use in therapeutic, diagnostic, or research applications. A kit may include one or more containers housing any of the vectors disclosed herein and instructions for use.

[0202] The kit may be designed to facilitate use of the methods described herein by researchers and can take many forms. Each of the compositions of the kit, where applicable, may be provided in liquid form (e.g., in solution), or in solid form, (e.g., a dry powder). In certain cases, some of the compositions may be constitutable or otherwise processable (e.g., to an active form), for example, by the addition of a suitable solvent or other species (for example, water or a cell culture medium), which may or may not be provided with the kit. As used herein, "instructions" can define a component of instruction and / or promotion, and typically involve written instructions on or associated with packaging of the disclosure. Instructions also can include any oral or electronic instructions provided in any manner such that a user will clearly recognize that the instructions are to be associated with the kit, for example, audiovisual (e.g., videotape, DVD, etc.), Internet, and / or web-based communications, etc. The written instructions, in some embodiments, are in a form prescribed by a governmental agency regulating the manufacture, use, or sale of pharmaceuticals or biological products, which instructions can also reflect approval by the agency of manufacture, use, or sale for animal administration.EXAMPLES

[0203] The following examples are given for the purpose of illustrating various embodiments of the disclosure and are not meant to limit the present disclosure in any fashion. The present examples, along with the methods described herein, are presently representative of preferred embodiments, are exemplary, and are not intended as limitations on the scope of the disclosure. Changes therein and other uses which are encompassed within the spirit of the disclosure as defined by the scope of the claims will occur to those skilled in the art.Attorney Docket No. MTGNZ000100WO

[0204] CRISPR-associated Transposases (CAST) are a new family of Tn7-like and Tn5053- like transposon systems, which encode RNA guided, nuclease-dead CRISPR Cas proteins for targeted integration of large cargoes. The Cas 12k CAST system encompasses a CRISPR Casl2k effector and three transposition proteins: TnsB, TnsC and TniQ. In addition, the Cas 12k CAST system requires the small prokaryotic ribosomal protein subunit S 15 for efficient cargo integration, and the accessory protein ClpX enhances transposition efficiency for the type V-K CAST MG64-1 system (Liu, Goltsman et al. 2025). Disclosed here is an optimized MG64-1 system through protein engineering, donor and CAST plasmid optimization, and CAST protein expression, and demonstrate >50X improved integration efficiency in mammalian cells.Example 1 - MG64-1 and MG64-6 are efficient systems for targeted genomic integration in E. coli

[0205] Construction ofE. coli with engineered target

[0206] For testing of effector assisted integrase activity in bacterial cells, strain MGB0034 was constructed from BL21(DE3) E. coli cells. A target sequence of 5’- GTCGAGGCTTGCGACGTGGTGGCT-3’ was inserted into the lacZ locus immediately downstream of a 5’-AGTC-3’ PAM sequence. MGB0034 E. coli cells were then transformed with two plasmids: pHelper and pGuide. pHelper is an ampicillin-resistant plasmid that expresses the effector and the Tns proteins suite for either MG64-1 or MG64-6 Tns proteins and Cas 12k. In multi-loci experiments, the pHelper of MG64-1 also contained a single guide targeted to the engineered target sequence driven by the T7 promoter. pGuide is a chloramphenicol resistant plasmid that expresses the single guide RNA sequence for the engineered or endogenous target of interest driven by a J23119 promoter.

[0207] E. coli Transposition Experiments

[0208] A culture containing pHelper or pHelper and pGuide plasmids was grown to at 37°C until saturation, diluted at least 1 :10 into LB with appropriate antibiotics (OD < 0.2), and incubated at 37°C until OD of approximately 0.6. Cells from this growth stage were made chemically competent by washing four times in a lx volume of ice-cold 0. 1 M calcium chloride. Cells were then transformed with 100 ng pDonor, a plasmid bearing a kanamycin or tetracycline resistance marker flanked by left end (LE) and right end (RE) transposon motifs for integration. Heat shocked cells were then recovered for 2 hours on LB medium at 37°C before being plated on LB-agar-ampicillin-chloramphenicol-kanamycin with or without 0.02 mM IPTG, and incubated two days at 37°C. Plates were scraped into LB medium, and a cell pellet was collected by centrifugation at 14,000 rpm for 30 minutes. The pellet wasAttorney Docket No. MTGNZOOOIOOWO resuspended in ~1 mL of LB. Approximately 200 pL of this suspension was set aside for genomic DNA extraction, while the remainder was re-pelleted and stored at -20°C. Genomic DNA was extracted and quantified. Transposition experiments were performed and analyzed in triplicate.

[0209] E. coli NGS Transposition Efficiency, on / off target determination, co-integration analysis

[0210] Whole genome sequencing was performed by building DNA sequencing libraries from extracted E. coli genomic DNA. Whole genome sequencing was also performed using the rapid barcoding kit. The results of short read whole genome sequencing following transposition experiments were analyzed by custom python scripts identifying moving averages of an on-target integration normalized to coverage of the integration site. Briefly, filtered reads were aligned against the MGB0034 genome sequence, pHelper, pTCM and pDonor plasmids using BWA 36. Chimeric reads mapping to both pDonor TIR sequences and genomic sequences were filtered from the alignment and normalized to the number of reads spanning at least 20 bp on either side of the genomic breakpoint. Total efficiency of transposition was calculated by performing a weighted average of local breakpoints across relative transposition efficiency. Breakpoint analysis was performed at the 5’-TGTACA-3’ motif of the LE and RE and on-target relative transposition efficiency was summed for breakpoints across a 20bp on-target window between 50-70bp distance from the PAM. All other breakpoints outside the 20bp window were counted as off-target.

[0211] Co-integration events were detected by aligning the long reads against the MGB0034 genome sequence, pHelper, pTCM and pDonor plasmids. Chimeric reads were extracted from the alignment and coordinates of the DNA cargo within the pDonor plasmid were projected onto read segments that aligned to the pDonor plasmid using custom python scripts. Chimeric reads in which a single segment of the read aligned to the full length of the DNA cargo, flanked on both sides by read segments that aligned to the E. coli genome, were counted as single integration events. Chimeric reads in which a single segment of the read aligned to the full length of the DNA cargo, flanked on one side by a read segment that aligned to the E. coli genome and on the other side a read segment that aligned to the pDonor backbone, were counted as co-integration events. Due to the error profile of these long reads, aligned segments in chimeric reads shorter than 25 bp were ignored during this analysis.

[0212] Results

[0213] Previously, the reference S. hoffmanni type V-K CAST (ShCAST) system’s activity has been shown to result in a mixture of integration events when the donor is delivered as a circular plasmid. In addition to the expected integration of the transposon cargo, up to 80% ofAttorney Docket No. MTGNZ000100WO integrations have been shown to include two copies of the cargo along with the plasmid backbone in what are referred to as co-integration events. This occurs due to the absence of the transposase protein TnsA for second strand donor cleavage. Both MG64-1 and MG64-6 systems lack TnsA and, as expected, result in both single (20-30%) and cointegration (70-80%) events upon delivery of a circular plasmid donor (FIG. 1A). In addition, as has been observed when the ShCAST system is used for integration in E. coli, the forward orientation of integration is favored over the reverse orientation (FIG. IB).Example 2 - Nuclear localization tags and fusion domains activate CAST for integration in the mammalian cell environment

[0214] Methods

[0215] To test the functionality of the NLS constructs in a physiologically relevant environment, constructs cloned with active NLS -tagged CAST components were integrated into K562 cells using lentiviral transduction. Briefly, constructs cloned into lentiviral transfer plasmids were transfected into 293T cells with envelope and packaging plasmids, and virus containing supernatant was harvested from the media after 72 hr incubation. Media containing virus was then incubated with K562 or HEK293T cell lines with 8 pg / mL of polybrene for 72 hrs, and transfected cells were then selected for integration in bulk using Puromycin at 1 g / mL for 4 days. Cell lines undergoing selection were harvested at the end of 4 days, and differentially lysed for nuclear and cytoplasmic fractions. Subsequent fractions were then tested for transposition capability with a complementary set of in vitro expressed components.

[0216] 10 million cells were harvested and washed once with IxPBS pH7.4. Supernatant wash was aspirated completely to the cell pellet, and flash frozen at -80C for 16 hrs. After thawing on ice, cell pellet size was measured by mass, and appropriate extraction volumes of cell fractionation and nuclear extraction reagent was used to natively extract proteins in cell fractions. Briefly, cytoplasmic extraction reagent was used at 1:10 mass of cells to volume of extraction reagent. Cell suspension was mixed by vortexing and lysed with non-ionic detergent. Cells are then centrifuged at 16,000xg at 4°C for 5 minutes. Cytoplasmic extraction supernatant was then decanted and saved for in vitro testing. Nuclear extraction reagent was then added 1:2 original cell mass to nuclear extraction reagent, and incubated on ice for 1 hr on ice with intermittent vortexing. Nuclear suspension was then centrifuged at 16,000 x g for 10 minutes at 4°C and supernatant nuclear extract was decanted and tested for in vitro transposition activity. Using 4 pL of each cell and nuclear extract for each condition, the in vitro transposition reaction was performed with a complementary set of in vitro expressedAttorney Docket No. MTGNZOOOIOOWO proteins, IVT sgRNA, donor DNA, pTarget, and buffer. Evidence of transposition activity was assayed by PCR amplification of donor-target junctions.

[0217] Results

[0218] In order to confirm that CAST protein components are active after successful nuclear localization with NLS tag fusions, nuclear extracts were obtained for each protein component and tested for integration activity in vitro (FIG. 2A). NLS-tagged Casl2k and TniQ from the MG64-1 CAST system were not active despite proper localization to the nucleus (FIG. 2B, lanes 8 and 10). However, when tested with NLS-tagged Casl2k and TniQ fusions with and without sso7d, human Hl -core, and human HMGN1 domains, fusions activate CAST extracted from the nuclear environment (FIG. 2B, lanes 4 and 6). Results suggest that chromatin accessibility domains are necessary accessory components for integration activity in mammalian cells.Example 3 - Nuclear localization tags and fusion domains for mammalian cell cargo integration are system-specific

[0219] Methods

[0220] The ShCAST protein components were synthesized as direct replacement within the MG64-1 all-in-one plasmid. Briefly, the all-in-one Helper plasmid encodes a single promoter driving all five protein coding components, along with a cloned single guide RNA target tailored to the AAVS1 locus (FIG. 3A). An additional plasmid containing the host factor ClpX-NLS (SEQ ID NO: 1395) was used in the integration reaction. The ShCAST donor was designed with LE and RE, synthesized on the donor plasmid where binding of oJL1125 was maintained for distance from the LE breakpoint for transposition. The sgRNA for ShCAST was synthesized with a “UUUCGUU” sequence replacing the native “UUUUGUU” motif in the tracrRNA in order to prevent early termination of transcription due to the presence of four consecutive U bases. For each MG64-1 and ShCAST systems, the three plasmids were delivery by LT1 transfection reagent at a ratio of 12 pg pHelper : 6 pg pDonor: 1 pg pClpX on 2,500,000 HEK293T cells in a 10 cm petri dish. Transfected cells were recovered for 72 hours at 37°C, genomic DNA (gDNA) was extracted from the cells at 72 hour post-transfection, and then quantified by NGS using LE primers (oJLl 109 and oJLl 125).

[0221] CAST integration quantification by NGS

[0222] Target specific donor plasmids were made so that a short genomic fragment of the AAVS 1 locus was cloned into the pDonor plasmid at a distance of 152 bp away from the Left End 5 ’ sequence. This genomic fragment allows for the simultaneous amplification of the AAVS1 transposition product and the genomic sequence with the same PCR primer set.Attorney Docket No. MTGNZ000100WOGenomic and transposed LE forward sequences were amplified using 100 pL PCR reactions of polymerase for 25 cycles using 500 nM oJLl 109 and oJLl 125 primers. Primers included an adaptor sequence and a 5 bp diversity stub. Sequencing adapters were then amplified for 10 cycles onto the PCR product library after lx SPRI cleanup and sequenced with a 2 x 300 cycle V3 kit. Resulting NGS reads were analyzed using the 63 bp target 1 transposition sequence as the reference amplicon, the genomic fragment as the HDR amplicon, and the Left End 5 ’ sequence as the spacer within a 20 bp window. Resulting alignments were then filtered for reads passing “-amas 95 filter”. Unmodified reference sequences, NHEJ sequences, HDR sequences, and modified HDR sequences were then pulled from the editing profile.Transposition frequencies were calculated by summing reference aligned sequences and NHEJ sequences over the total amount of reads.

[0223] Results

[0224] ShCAST cognate components were synthesized with NLS fusions in the same orientation as those for MG64-1 , and chromodomains sso7d and Hl core were fused to the C- terminus of the ShCAST Casl2k and N-terminus of TniQ, respectively (SEQ ID NOs: 1394, 1398-1401, and 1406-1409). sgRNA optimization (SEQ ID NOs: 1402-1403 and 1410-1411) was applied for expression in human cells, and cargoes for with MG64-1 LE and RE (SEQ ID NOs: 1404-1405), or for ShCAST (SEQ ID NOs: 1396-1397) were designed using the full- length LE and RE elements. Given that ShCAST recognizes a 5’ NGTN PAM, the same AAVS1-T1 targeting spacer was used that was validated with MG64-1 (SEQ ID NO: 1387). When tested under the same experimental conditions with ClpX co-expression, MG64-1 CAST showed 2% targeted integration as measured by NGS, while for ShCAST, no reads could be identified that would indicate target-specific integration (FIG. 3B). Results suggest that type V-K CAST optimizations for mammalian cell cargo integration are system-specific.Example 4 - Protein engineering improves integration efficiency of MG64-1 in HEK293T cells

[0225] Results: engineered TnsB and Casl2k ASR variants

[0226] WT type V-K CAST MG64-1 may be capable of programmable integration in human cells when chromodomains were fused to the Casl2k effector and TniQ transposition component. TnsB was engineered by rationally selecting single mutations, as well as by designing a mutagenesis library of several surface sites identified from the diversity of residues at these sites observed in TnsB homologs from other organisms. In addition, Casl2k variants were designed through ancestral sequence reconstruction (ASR). Results indicate that combining Casl2k ASR and TnsB variants may improve integration efficiency of a 4.7 kbAttorney Docket No. MTGNZ000100WO donor gene at AAVS1 in human HEK293T cells by >50X (FIG. 4A, Casl2k ASR + TnsB v3 variant).

[0227] Results: delivery of donor and CAST components in non-replicating plasmids

[0228] To determine if improvements observed from engineered variants in HEK293T cells would translate to other cell types (e.g. due to differences in DNA repair, lack of DNA replication, etc.), the donor and CAST components were delivered in replicating and nonreplicating plasmids. It was demonstrated that integration efficiency is not dependent on replicating plasmids in HEK293T, with variants and donor from non-replicating plasmids significantly outperforming those from replicating plasmids (FIG. 5, bottom).

[0229] In addition, to determine the order in which the CAST components are expressed off of the plasmid, a library was designed that measured all possible linear arrangements of the four CAST components (FIG. 5, top). It was observed that plasmids that express TnsC before other components, engineered TnsB and Casl2k ASR variants achieve higher integration efficiencies compared to control plasmids with WT variants, where TnsC is the last component (Fig. 2, bottom). Overall, combination of improvements consistently results in >4% integration efficiency as measured by NGS of the target-to-left end junction.Example 5 - Combination of improved MG64-1 variants CAST variants delivered by mRNA results in >1 % integration efficiency at the ALB locus in primary mouse hepatocytes

[0230] Results: combination of improvements translate to hepatocyte cells

[0231] To determine if improved efficiencies observed in HEK293T cells can also be observed in other cell types, MG64-1 CAST variants were tested for integration of a donor in Hep3B immortalized hepatocytes. 90k Hep3B cells were seeded in a 24-well plate and transfected the next day with Lipofectamine 3000. Cells were either transfected with 333ng pHelper: 800ng pDonor or 3ug mRNA pHelper: 800ng pDonor. gDNA was extracted after 72h. Results confirmed that Improvements observed in HEK293T are translatable to immortalized hepatocytes, with observed integration efficiencies >0.4% (FIG. 4B).

[0232] In addition, to determine if improved variants were active for programmable integration in primary cells, MG64-1 CAST variants were tested in primary mouse hepatocytes (PMH). 100k PMH were seeded in a 24-well plate 24h prior to transfection. An all-in-one (AIO) mRNA encoding improved CAST protein variants was delivered, with or without ClpX mRNA and donor DNA plasmid containing factor VIII (gene of interest, GOI) with Lipofectamine 2000 or a combination of Lipofectamine 2000 and messengerMAX, simultaneously or in a 6- hour interval lipofection reaction (staggered), at a 3:1 or 6.8:1.8 dose ratio. Results indicate that simultaneous delivery of 3 : 1 AIO to donor dose achieves the highest integration efficiencyAttorney Docket No. MTGNZ000100WO in PMH at nearly 2% efficiency, as measured by NGS amplification of the target-to-left end junction product (FIG. 6). ClpX did not improve editing with CAST variants in PMH, indicating that the enhancer is no longer needed for targeted integration in cells.Example 6 - Protein engineering improves integration efficiency of MG63-1 in HEK293T cells

[0233] Methods: TnsB library design and screening in human cells

[0234] A deep mutational scanning library of 15 sites located in the second shell of the catalytic center of MG64-1 TnsB was designed. In addition, a combinatorial library was designed of up to 10 sites in TnsB selected based on naturally occurring diversities identified in a multiple sequence alignment of 260 TnsB homologous sequences (SEQ ID NO. 1425-1429). Human codon optimized sequences for the variants were cloned into the AIO plasmid encoding all other CAST protein components for screening in human cells.

[0235] Methods: HEK293T cell transfection and screening of variants

[0236] Screening for integration activity with newly-designed TnsB variants was done using three plasmids delivered via lipofection, as previously described in Examples 1-5 above. Briefly, fragments for designed library variants were cloned into an all-in-one (AIO) pHelper plasmid featuring a single promoter driving all protein coding components (FIG. 7A). A second plasmid contained the LE and RE (TIR) flanking a donor cargo that included a chimeric apolipoprotein E human al anti-trypsin (ApoE-hAAT) promoter (SEQ ID NO. 1421) driving a full Factor IX gene (SEQ ID NO. 1422), and either one or two copies of the single guide RNA encoding the spacer for the AAVS1 target 5 locus and an AAVS1 target 5 PBS (SEQ ID NO. 1414). In some experiments, a third replicative plasmid contained the ClpX-NLS enhancer (FIG. 7A). The three- plasmid system was combined at ratios of 0.165 ug pHelper, 0.165 ug of pDonor, and 0.025 ug of pClpX and delivered to 25,000 HEK293T cells in a 96 well plate with Lipofectamine-2000 transfection reagent. Transfected cells were recovered for 72 hours at 37°C, gDNA was extracted from the cells, and editing of alleles was calculated for integration events by amplifying and sequencing via NGS the target-to-LE junctions.

[0237] Results: engineered TnsB variants improve integration efficiency to >4%

[0238] The MG64-1 TnsB was engineered by rationally selecting single residue mutations for 15 sites in the second shell of the catalytic center, as well as by designing a combinatorial mutagenesis library of high frequency natural diversities in sites with medium to high mutational entropy identified from a multiple sequence alignment of TnsB homologs. Several TnsB single amino acid substitution variants were found to have multiple fold improvement over the WT control, including R347A, R347W, R347I, I228M (SEQ ID NOs. 1425, 1427-1429) (FIG. 8A). For some engineered variants, integration efficiency was higher when the ClpX plasmid was notAttorney Docket No. MTGNZOOOIOOWO added to the reaction (FIG. 8B). For TnsB combinatorial mutations, tetramutant and hexamutant variants (SEQ ID NOs. 1430-1431) achieved on average 3% integration efficiency with addition of the ClpX plasmid (Fig. 9).

[0239] While preferred embodiments of the present disclosure have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the disclosure be limited by the specific examples provided within the specification. While the disclosure has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the disclosure. Furthermore, it shall be understood that all aspects of the disclosure are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the disclosure described herein may be employed in practicing the disclosure. It is therefore contemplated that the disclosure shall also cover any such alternatives, modifications, variations, or equivalents. It is intended that the following claims define the scope of the disclosure and that methods and structures within the scope of these claims and their equivalents be covered thereby.Attorney Docket No. MTGNNZ00100PREFERRED EMBODIMENTS

[0240] (1) In one preferred embodiment, there is a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid comprising: a) a Cas effector complex and a Tn7 type transposase complex that binds the Cas effector complex, wherein the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence having at least 70% sequence identity to any one of SEQ ID NOs: 1406-1409; and b) a double-stranded nucleic acid that interacts with the Tn7 type transposase complex and comprises the cargo nucleotide sequence.

[0241] (2) In a further preferred embodiment, the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1406-1409.

[0242] (3) In a further preferred embodiment, the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1406-1409.

[0243] (4) In a further preferred embodiment, the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence of any one of SEQ ID NOs: 1406-1409.

[0244] (5) In a further preferred embodiment, the Cas effector complex binds non-covalently to the Tn7 type transposase complex.

[0245] (6) In a further preferred embodiment, the Cas effector complex is covalently linked to the Tn7 type transposase complex.

[0246] (7) In a further preferred embodiment, the Cas effector complex is fused to the Tn7 type transposase complex.

[0247] (8) In a further preferred embodiment, the cargo nucleotide sequence is flanked by a lefthand transposase recognition sequence and a right-hand transposase recognition sequence recognized by the Tn7 type transposase complex.

[0248] (9) In a further preferred embodiment, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least 80% identity to SEQ ID NO: 1404.

[0249] (10) In a further preferred embodiment, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least 80% identity to SEQ ID NO: 1405.

[0250] (11) In a further preferred embodiment, the engineered guide polynucleotide comprises a nucleotide sequence comprising at least about 46-80 consecutive nucleotides having at least 80% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411.Attorney Docket No. MTGNNZ00100

[0251] (12) In a further preferred embodiment, the engineered guide polynucleotide comprises a nucleotide sequence comprising having at least 80% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411.

[0252] (13) In one preferred embodiment, there is a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid comprising: a) a Cas effector complex and a Tn7 type transposase complex that binds the Cas effector complex, wherein the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence having at least 70% sequence identity to any one of SEQ ID NOs: 1406-1409; and b) a double-stranded nucleic acid that interacts with the Tn7 type transposase complex and comprising in 5’ to 3’ order: i) a left-hand transposase recognition sequence comprising a nucleotide sequence having at least 80% sequence identity to SEQ ID NO: 1404; ii) the cargo nucleotide sequence; and iii) a right-hand transposase recognition sequence comprising a nucleotide sequence having at least 80% identity to SEQ ID NO: 1405.

[0253] (14) In a further preferred embodiment, the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1406-1409.

[0254] (15) In a further preferred embodiment, the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1406-1409.

[0255] (16) In a further preferred embodiment, the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence of any one of SEQ ID NOs: 1406-1409.

[0256] (17) In a further preferred embodiment, the Cas effector complex binds non-covalently to the Tn7 type transposase complex.

[0257] (18) In a further preferred embodiment, the Cas effector complex is covalently linked to the Tn7 type transposase complex.

[0258] (19) In a further preferred embodiment, the Cas effector complex is fused to the Tn7 type transposase complex.

[0259] (20) In a further preferred embodiment, the cargo nucleotide sequence is flanked by a left-hand transposase recognition sequence and a right-hand transposase recognition sequence recognized by the Tn7 type transposase complex.

[0260] (21) In a further preferred embodiment, the left-hand transposase recognition sequence comprises a nucleotide sequence having at least 80% identity to SEQ ID NO: 1404.Attorney Docket No. MTGNNZ00100

[0261] (22) In a further preferred embodiment, the right-hand transposase recognition sequence comprises a nucleotide sequence having at least 80% identity to SEQ ID NO: 1405.

[0262] (23) In a further preferred embodiment, the engineered guide polynucleotide comprises a nucleotide sequence comprising at least about 46-80 consecutive nucleotides having at least 80% identity to SEQ ID NO: 1410 or SEQ ID NO: 141 1.

[0263] (24) In a further preferred embodiment, the engineered guide polynucleotide comprises a nucleotide sequence comprising having at least 80% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411.

[0264] (25) In one preferred embodiment, there is a method for transposing a cargo nucleotide sequence into a target nucleic acid site comprising introducing the system of any one of the preferred embodiments disclosed herein to a cell.

[0265] (26) In one preferred embodiment, there is a cell comprising the system of any one of the preferred embodiments disclosed herein.

[0266] (27) In a further preferred embodiment, the cell is a eukaryotic cell.

[0267] (28) In a further preferred embodiment, the cell is a mammalian cell.

[0268] (29) In a further preferred embodiment, the cell is an immortalized cell.

[0269] (30) In a further preferred embodiment, the cell is an insect cell.

[0270] (31) In a further preferred embodiment, the cell is a yeast cell.

[0271] (32) In a further preferred embodiment, the cell is a plant cell.

[0272] (33) In a further preferred embodiment, the cell is a fungal cell.

[0273] (34) In a further preferred embodiment, the cell is a prokaryotic cell.

[0274] (35) In a further preferred embodiment, the cell is an A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Veto, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cell, HT1080, HepG2, Huh7, K562, primary cell, or a derivative thereof.

[0275] (36) In a further preferred embodiment, the cell is an engineered cell.

[0276] (37) In a further preferred embodiment, the cell is a stable cell.

Claims

Attorney Docket No. MTGNNZ00100CLAIMSWHAT IS CLAIMED IS:

1. A system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid comprising: a) a Cas effector complex and a Tn7 type transposase complex that binds the Cas effector complex, wherein the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence having at least 70% sequence identity to any one of SEQ ID NOs: 1406-1409; and b) a double-stranded nucleic acid that interacts with the Tn7 type transposase complex and comprises the cargo nucleotide sequence.

2. The system of claim 1, wherein the Cas effector complex: a) binds non-covalently to the Tn7 type transposase complex; b) is covalently linked to the Tn7 type transposase complex; or c) is fused to the Tn7 type transposase complex.

3. The system of claims 1 and 2, wherein the cargo nucleotide sequence is flanked by a lefthand transposase recognition sequence and a right-hand transposase recognition sequence recognized by the Tn7 type transposase complex.

4. The system of claim 3, wherein the left-hand transposase recognition sequence comprises a nucleotide sequence having at least 80% identity to SEQ ID NO: 1404.

5. The system of claim 3, wherein the right-hand transposase recognition sequence comprises a nucleotide sequence having at least 80% identity to SEQ ID NO: 1405.

6. The system of any one of claims 1-5, wherein the engineered guide polynucleotide comprises a nucleotide sequence comprising at least about 46-80 consecutive nucleotides having at least 80% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411.

7. A system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid comprising: a) a Cas effector complex and a Tn7 type transposase complex that binds the Cas effector complex, wherein the Cas effector complex and the Tn7 type transposase complex comprises a nucleotide sequence having at least 70% sequence identity to any one of SEQ ID NOs: 1406-1409; and b) a double-stranded nucleic acid that interacts with the Tn7 type transposase complex and comprising in 5’ to 3’ order: i) a left-hand transposase recognition sequence comprising a nucleotide sequence having at least 80% sequence identity to SEQ ID NO: 1404; ii) the cargo nucleotide sequence; andAttorney Docket No. MTGNNZ00100 iii) a right-hand transposase recognition sequence comprising a nucleotide sequence having at least 80% identity to SEQ ID NO: 1405.

8. The system of claim 7, wherein the Cas effector complex: a) binds non-covalently to the Tn7 type transposase complex; b) is covalently linked to the Tn7 type transposase complex; or c) is fused to the Tn7 type transposase complex.

9. The system of claims 7 and 8, wherein the cargo nucleotide sequence is flanked by a lefthand transposase recognition sequence and a right-hand transposase recognition sequence recognized by the Tn7 type transposase complex.

10. The system of claim 9, wherein the left-hand transposase recognition sequence comprises a nucleotide sequence having at least 80% identity to SEQ ID NO: 1404.

11. The system of claim 9, wherein the right-hand transposase recognition sequence comprises a nucleotide sequence having at least 80% identity to SEQ ID NO: 1405.

12. The system of any one of claims 7-11, wherein the engineered guide polynucleotide comprises a nucleotide sequence comprising at least about 46-80 consecutive nucleotides having at least 80% identity to SEQ ID NO: 1410 or SEQ ID NO: 1411.

13. A method for transposing a cargo nucleotide sequence into a target nucleic acid site comprising introducing the system of any one of claims 1-12 to a cell.

14. A cell comprising the system of any one of claims 1-12.

15. The cell of claim 14, wherein the cell is a eukaryotic cell, a mammalian cell, an immortalized cell, an insect cell, a yeast cell, a plant cell, a fungal cell, or a prokaryotic cell.

16. The cell of claim 14, wherein the cell is an A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cell, HT1080, HepG2, Huh7, K562, primary cell, or a derivative thereof.

17. The cell of claim 14, wherein the cell is an engineered cell.

18. The cell of claim 14, wherein the cell is a stable cell.

Citation Information

Patent Citations

  • Novel crispr-associated transposon systems and components

    US20200291395A1

  • Fusion proteins

    WO2023164592A2