Engineered programmable DNA transposases and engineered bridge RNA systems for nucleic acid manipulation

The IS110 family transposase and bridgeRNA system addresses the challenges of reprogramming DNA recombination systems by enabling efficient and precise DNA manipulation, including insertion, inversion, and excision, with enhanced recombination efficiency and simplified delivery.

WO2025160203A1PCT designated stage expired Publication Date: 2025-07-31ARC RES INST +2
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/012632
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-17
Filing Date
2025-01-22
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Current DNA sequence manipulation methods, such as the Cre-loxP system and RNA-guided programmable nucleases, face challenges in reprogramming DNA sequence recombination systems without significant protein engineering efforts, and delivery to cells is complex, especially for efficient inversion or excision of DNA sequences.

Method used

A recombinant nucleic acid editing system utilizing an IS110 family transposase and a bridgeRNA with handshake guide sequences to facilitate programmable DNA insertion, inversion, and excision, enhancing recombination efficiency and specificity.

Benefits of technology

The system enables precise and efficient manipulation of DNA sequences, including insertion, inversion, and excision, with improved recombination efficiency and reduced complexity in delivery to cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025012632_31072025_PF_FP_ABST
    Figure US2025012632_31072025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to novel systems for nucleic acid engineering utilizing components of IS110 family transposons including a modified bridgeRNA. The bridgeRNA, which targets donor and target sequence sites or multiple donor site sequences for polynucleotide recombination reactions, is modified within the handshake guide (HSGs) sequences, one or more scaffold sequences, or both the HSG sequences and one or more scaffold sequences. In certain aspects, the application relates to the utilization and reprogramming of the modified bridgeRNA to direct IS110 transposases to integrate sequences at predetermined sites. Programmable insertion, excisive recombination, and / or inversion enables integration or transposition of any polynucleotide sequence encoding a donor site or target site recognized by the IS110 transposase into any other polynucleotide sequence containing a target site sequence or donor site sequence, respectively, or between two donor site sequences using an IS110 transposase and modified bridgeRNA.
Need to check novelty before this filing date? Find Prior Art

Description

ENGINEERED PROGRAMMABLE DNA TRANSPOSASES AND ENGINEERED BRIDGE RNA SYSTEMS FOR NUCLEIC ACID MANIPULATION

[0001] This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 623,721 filed January 22, 2024, and U.S. Provisional Patent Application No. 63 / 649,314 filed May 17, 2024, the contents of each of which are hereby incorporated by reference in their entireties.

[0002] All patents, patent applications and publications cited herein are hereby incorporated by reference in their entirety. The disclosures of these publications in their entireties are hereby incorporated by reference into this application.

[0003] This patent disclosure contains material that is subject to copyright protection. The copyright owner has no objection to the facsimile reproduction by anyone of the patent document or the patent disclosure as it appears in the U.S. Patent and Trademark Office patent file or records, but otherwise reserves any and all copyright rights.SEQUENCE LISTING

[0004] The instant application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on January 21, 2025, is named 2220476-00125W01_SL.xml and is 58,323,000 bytes in size.BACKGROUND OF THE INVENTION

[0005] DNA sequence manipulation, namely insertion, inversion, and excision of DNA, is a foundational capability underpinning synthetic biology, genome engineering, cell engineering, and genetic medicine. Such reactions have been utilized for decades, such as the Cre-loxP system, which relies on the Cre protein's specificity for pre-engineered loxP DNA sequences. The orientation and location of these DNA sequences relative to one another are used to achieve insertion, inversion, or excision of long DNA sequences. More modem methods, especially for DNA insertion, often utilize random, semi-random, or nonprogrammable site-specific enzymes such as recombinases and transposases for manipulation of DNA sequences, relying on enzyme-encoded specificity for target sequences. RNA-guided programmable nucleases such as Cas9, have been utilized to excise DNA by cutting at morethan one site, albeit relying on DNA repair of the resulting double stranded break. For insertion and inversion, the amino acid-encoded DNA specificities of these systems are difficult to reprogram without significant protein engineering efforts, in particular through the tethering of a separate programmable targeting domain, such as an RNA-guided CRISPR effector. Other methods utilize synthetic or naturally occurring reprogrammable transposition systems that require the assembly of multiple protein and nucleic acid subunits to achieve DNA integration, yet delivery of these systems to cells can be complex and they require extensive engineering to achieve efficient inversion or excision of DNA sequences. While DNA sequence recombination systems that are reprogrammable have been described, there is a need for further improved DNA sequence recombination systems.SUMMARY OF THE INVENTION

[0006] It is understood that any of the embodiments described below can be combined in any desired way, and that any embodiment or combination of embodiments can be applied to each of the aspects described below, unless the context indicates otherwise.

[0007] In certain aspects, the invention provides a recombinant nucleic acid editing system comprising: a) an IS 110 family transposase, or a nucleic acid comprising a sequence encoding the IS110 family transposase; and b) a nucleic acid comprising a sequence encoding a bridgeRNA, wherein the bridgeRNA comprises one or more handshake guide (HSG) sequences.

[0008] In some embodiments, the nucleic acid comprising a sequence encoding a bridgeRNA comprises a left end (LE) sequence of a transposon that encodes the IS110 family transposase. In some embodiments, the nucleic acid comprising a sequence encoding a bridgeRNA comprises a right end (RE) sequence of a transposon that encodes the IS110 family transposase.

[0009] In some embodiments, the recombinant nucleic acid editing system further comprises a nucleic acid comprising a RE sequence, an anti-HSG donor site sequence, and a LE sequence or a RE sequence, an anti-HSG donor site sequence, a core sequence, and a LE sequence of the IS110 element that encodes the IS110 family transposase. In some embodiments, the recombinant nucleic acid editing system further comprises a nucleic acid comprising a right flank (RF) sequence, an anti-HSG target site sequence, and a left flank (LF) sequence or a RF sequence, an anti-HSG target site sequence, a core sequence, and a LF sequence of the target site sequence for the IS110 family transposase. In some embodiments,the nucleic acid comprising the RE sequence, the anti-HSG donor site sequence, and the LE sequence or the RE sequence, the anti-HSG donor site sequence, the core sequence, and the LE sequence further comprises a nucleic acid sequence for insertion into a target site sequence. In some embodiments, the target site sequence comprises a RF sequence, an anti- HSG target site sequence, and a LF sequence or a RF sequence, an anti-HSG target site sequence, a core sequence, and a LF sequence for the IS110 family transposase. In some embodiments, the nucleic acid comprising the RF sequence, the anti-HSG target site sequence and the LF sequence or the RF sequence, the anti-HSG target site sequence, the core sequence, and the LF sequence further comprises a nucleic acid sequence for insertion into a donor site sequence. In some embodiments, the donor site sequence comprises a RE sequence, an anti-HSG donor site sequence, and a LE sequence or a RE sequence, an anti- HSG donor site sequence, a core sequence, and a LE sequence of the IS110 element that encodes the IS110 family transposase. In same embodiments, said target and / or said donor sites are chosen such that one or more of the HSGs of said bridgeRNA base pair to enhance recombination.

[0010] In some embodiments, the bridgeRNA comprises a nucleotide sequence at least 50% identical to a bridgeRNA sequence of SEQ ID NOS: 1-348 or SEQ ID NOS: 349-10175. In some embodiments, said RE sequence comprises a RE sequence of SEQ ID NOS: 1-348, 30354-30529, 349-10175 or 30530-40356, said LE sequence comprises a LE sequence of SEQ ID NOS: 1-348, 30354-30529, 349-10175 or 30530-40356, and / or said core sequence comprises a core sequence of SEQ ID NOS: 1-348, 30354-30529, 349-10175 or 30530- 40356.

[0011] In certain aspects, the invention provides a recombinant nucleic acid editing system comprising: a) an IS 110 family transposase, or a nucleic acid comprising a sequence encoding the IS110 family transposase; and b) a bridgeRNA, or a nucleic acid comprising a sequence encoding the bridgeRNA, the bridgeRNA comprising at least one stem-loop structure and further comprising at least one internal loop comprising a first nucleotide sequence that is complementary to a first target site sequence of a target DNA, a second nucleotide sequence that is complementary to a second target site sequence which is on the opposite strand of the target DNA to the first target site sequence, and a handshake guide nucleotide sequence that is complementary to a donor site sequence, and wherein the bridgeRNA is capable of forming a complex with the IS110 family transposase.

[0012] In some embodiments, the bridgeRNA further comprises a third nucleotide sequencethat is complementary to a first donor site sequence of a donor DNA, a fourth nucleotide sequence that is complementary to a second donor site sequence which is on the opposite strand of the donor DNA to the first donor site sequence, and a handshake guide nucleotide sequence that is complementary to a target site sequence. In some embodiments, the third nucleotide sequence the fourth nucleotide sequence, and handshake guide nucleotide sequence that is complementary to a target site sequence are on a second internal loop.

[0013] In some embodiments, the IS110 family transposase comprises a RuvC-like DEDD catalytic domain and a transposase domain. In some embodiments, the IS110 family transposase further comprises a linker domain between the RuvC-like DEDD catalytic domain and transposase domain. In some embodiments, the linker domain comprises a coiled-coil linker domain. In some embodiments, the RuvC-like DEDD catalytic domain comprises an amino acid sequence at least 50% identical to a RuvC-like DEDD catalytic domain sequence of SEQ ID NOs: 10176-10523 or 10524-20350. In some embodiments, the IS110 family transposase comprises a RuvC-like DEDD catalytic domain that forms a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621. In some embodiments, the IS110 family transposase RuvC-like DEDD catalytic domain comprises a tertiary structure similar to a tertiary structure of the RuvC-like DEDD catalytic domain of IS621 if the template modeling score (TM-score) for the RuvC-like DEDD catalytic domain of the IS110 family transposase is 0.5 or higher. In some embodiments, the RuvC-like DEDD catalytic domain comprises an amino acid sequence at least 15% identical to a RuvC-like DEDD catalytic domain sequence of SEQ ID NOs: 10176-10523 or 10524-20350.

[0014] In some embodiments, the transposase domain comprises an amino acid sequence at least 50% identical to a transposase domain sequence of SEQ ID NOs: 10176-10523 or 10524-20350. In some embodiments, the IS110 family transposase comprises a transposase domain that forms a similar tertiary structure to the transposase domain of IS621. In some embodiments, the IS110 family transposase domain comprises a tertiary structure similar to a tertiary structure of the transposase domain of IS621 if the template modeling score (TM- score) for the transposase domain of the IS110 family transposase is 0.5 or higher. In some embodiments, the transposase domain comprises an amino acid sequence at least 15% identical to a transposase domain sequence of SEQ ID NOs: 10176-10523 or 10524-20350.

[0015] In some embodiments, the IS110 family transposase comprises an amino acid sequence at least 50% identical to SEQ ID NOs: 10176-10523 or 10524-20350. In some embodiments, the IS110 family transposase further comprises a tertiary structure similar to atertiary structure of IS621. In some embodiments, the IS110 family transposase comprises a tertiary structure similar to a tertiary structure of IS621 if the template modeling score (TM- score) for the transposase is 0.5 or higher. In some embodiments, the transposase domain comprises an amino acid sequence at least 15% identical to a transposase domain sequence of SEQ ID NOs: 10176-10523 or 10524-20350.

[0016] In some embodiments, the IS110 family transposase is an IS 110 group transposase. In some embodiments, the IS110 family transposase is an IS1111 group transposase. In some embodiments, the IS1111 group transposase is IS1111 A or IS1111 229727. In some embodiments, wherein the IS110 group transposase is IS621, ISPal 1, IsPa29, ISMmgl, ISPfll, ISMae40, ISStma6, ISAzs32, ISMex9, ISCARN28, ISAarl6, ISCps7, ISPpu9, ISRel9, ISEsa2, ISMma5, IS900, or ISHne5. In some embodiments, the IS110 group transposase comprises an amino acid sequence at least 50% identical to IS621 (SEQ ID NO: 10176). In some embodiments, the IS110 group transposase comprises an amino acid sequence at least 50% identical to IS621_23122 (SEQ ID NO: 10181).

[0017] In some embodiments, the bridgeRNA comprises at least two stem-loop structures comprising a first stem-loop and a second stem-loop where the first stem-loop is 5' to the second stem-loop and wherein the first stem-loop comprises a target binding loop and the second stem-loop comprises a donor binding loop. In some embodiments, the bridgeRNA further comprises a third stem-loop structure 5' of the first stem-loop. In some embodiments, the stem of the first stem-loop is 5 to 35 nucleotides long and the loop is 3-10 nucleotides long, the target binding loop is 5 to 20 nucleotides long, the stem of the second stem-loop is 5 to 35 nucleotides long and the loop is 3-10 nucleotides long, and the donor binding loop is 5 to 20 nucleotides long. In some embodiments, the stem of the second stem-loop structure comprises 1 to 4 loops or bubbles that are each 1 to 10 nucleotides long.

[0018] In some embodiments, the bridgeRNA comprises a nucleotide sequence comprising any of the 5' to 3' sequences provided in Figure 19, wherein “n” represents any nucleotide, “R” represents an A or G nucleotide, and “Y” represents a C or U nucleotide. In some embodiments, the bridgeRNA comprises a 5' to 3' secondary structure provided in the first row, second row, third row, or fourth row of secondary structure for said sequence provided in Figure 19, wherein matching parentheses “(“ and “)” indicate base-paired nucleotides, and indicates unpaired bases.

[0019] In some embodiments, the bridgeRNA comprises a stem-loop structure as depicted in Figure 2D, Figure 1 IB, or Figure 13. In some embodiments, the target bindingloop of the bridgeRNA comprises: a left-target guide (LTG) comprising, in the 5' to 3' direction, a nucleotide sequence complementary to first strand of a target site sequence wherein the 3' end of the LTG is complementary to at least one of the nucleotides of a core sequence on the first strand of the target site sequence; a right-target guide (RTG) comprising, in the 5' to 3' direction, a nucleotide sequence that is a reverse complement to an opposite strand of the first strand of the target site sequence wherein the 3' end of the RTG is reverse complementary to at least one of the nucleotides of the core sequence on the opposite strand of the first strand of the target site sequence, a handshake guide (HSG-TBL) comprising, in the 5' to 3' direction, a nucleotide sequence that is partially or fully, complementary to the opposite strand of the first strand of the donor site sequence (anti-HSG- donor site sequence) and wherein optionally the HSG-TBL is located 3' to the 3' end of the RTG or wherein the HSG-TBL is located 3' of the 3' end of the RTG and wherein the HSG- TBL is separated from the RTG by 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides; wherein the target site sequence is a polynucleotide sequence; and / or wherein the donor binding loop of the bridgeRNA comprises: a left-donor guide (LDG) comprising, in the 5' to 3' direction, a nucleotide sequence complementary to a first strand of a donor site sequence wherein the 3' end of the LDG complementary to at least one of the nucleotides of a core sequence on the first strand of the donor site sequence; a right-donor guide (RDG) comprising, in the 5' to 3' direction, a nucleotide sequence that is a reverse complement to an opposite strand of the first strand of the donor site sequence wherein the 3' end of the RDG is reverse complementary to at least one of the nucleotides of the core sequence on the opposite strand to the first strand of the donor site sequence, a handshake guide (HSG-DBL), comprising, in the 5' to 3' direction, a nucleotide sequence that is partially or fully complementary to the opposite strand of the first strand of the donor site sequence (anti-HSG- donor site sequence) and wherein the HSG-DBL is optionally located 3' to the 3' end of the RDG or wherein the HSG-DBL is located 3' of the 3' end of the RDG and wherein the HSG- DBL is separated from the RDG by 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides; wherein the donor site sequence is a polynucleotide sequence, and the core sequence of the donor site sequence is the same as the core sequence of the target site sequence. In some embodiments the HSG-TBL is partially complementary or fully complementary to one or more nucleotides immediately 5' to the core sequence on the opposite strand of the donor site sequence. In some embodiments the HSG-TBL is partially complementary or fully complementary to one or more nucleotides located 5' to the core sequence on the oppositestrand of the donor site sequence and wherein the one or more nucleotides are separated from the core sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides. In some embodiments the HSG-DBL is partially complementary or fully complementary to one or more nucleotides immediately 5' to the core sequence on the opposite strand of the target site sequence. In some embodiments the HSG-DBL is partially complementary or fully complementary to one or more nucleotides located 5' to the core sequence on the opposite strand of the target site sequence and wherein the one or more nucleotides are separated from the core sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides.

[0020] In some embodiments, the target binding loop of the bridgeRNA comprises: a lefttarget guide (LTG) comprising, in the 5' to 3' direction, a nucleotide sequence is reverse complementary to an opposite strand to a first strand of a target site sequence wherein the 3' end of the LTG is complementary to at least one of the nucleotides of a core sequence on the opposite strand to the first strand of the target site sequence; a right-target guide (RTG) comprising, in the 5' to 3' direction, a nucleotide sequence that is a complementary to the first strand of the target site sequence wherein the 3' end of the RTG is complementary to at least one of the nucleotides of the core sequence on the first strand of the target site sequence; and a handshake guide (HSG-TBL) comprising, in the 5' to 3' direction, a nucleotide sequence that is partially or fully complementary to the first strand of the donor site sequence (the anti- HSG-donor site sequence) and wherein optionally the HSG-TBL is located 3' to the 3' end of the RTG or wherein the HSG-TBL is located 3' of the 3' end of the RTG and wherein the HSG-TBL is separated from the RTG by 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides; wherein the target site sequence is a polynucleotide sequence; and / or wherein the donor binding loop of the bridgeRNA comprises: a left-donor guide (LDG) comprising, in the 5' to 3' direction, a nucleotide sequence complementary to a first strand of a donor site sequence wherein the 3' end of the LDG complementary to at least one of the nucleotides of a core sequence on the first strand of the donor site sequence; a right-donor guide (RDG) comprising, in the 5' to 3' direction, a nucleotide sequence that is a reverse complement to an opposite strand of the first strand of the donor site sequence wherein the 3' end of the RDG is reverse complementary to at least one of the nucleotides of the core sequence on the opposite strand to the first strand of the donor site sequence, and a handshake guide (HSG-DBL), comprising, in the 5' to 3' direction, a nucleotide sequence that is not, or is not fully, complementary to the first strand of the donor site sequence (the anti-HSG-donor site sequence) and wherein optionally the HSG-DBL is located 3' to the 3' end of the RDG orwherein the HSG-DBL is located 3' of the 3' end of the RDG and wherein the HSG-DBL is separated from the RDG by 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides; wherein the donor site sequence is a polynucleotide sequence, and the core sequence of the donor site sequence is the same as the core sequence of the target site sequence. In some embodiments the HSG-TBL is partially complementary or fully complementary to one or more nucleotides immediately 5' to the core sequence on the first strand of the donor site sequence. In some embodiments the HSG-TBL is partially complementary or fully complementary to one or more nucleotides located 5' to the core sequence on the opposite strand of the donor site sequence and wherein the one or more nucleotides are separated from the core sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides. In some embodiments the HSG-DBL is partially complementary or fully complementary to one or more nucleotides immediately 5' to the core sequence on the first strand of the target site sequence. In some embodiments the HSG-DBL is partially complementary or fully complementary to one or more nucleotides located 5' to the core sequence on the opposite strand of the target site sequence and wherein the one or more nucleotides are separated from the core sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides.

[0021] In some embodiments, the target site sequence comprises sequence X1X2X3X4X5X6X7X8X9X10X11X12X13X14 where X is any nucleotide, and XsX9 are the core, one or more of X1X2X3X4X5X6X7 are the anti-HSG-target site sequence, optionally wherein Xe and X7 are the anti-HSG-target site sequence, and one or more of X12, X13, and X14 are optionally part of the target site sequence. In some embodiments, the bridgeRNA encodes an LTG in the 5' to 3' direction X1X2X3X4X5X6X7X8 or X1X2X3X4X5X6X7X8X9, an RTG in the 5' to 3' direction Y14Y13Y12Y11Y10Y9Y8 or Y14Y13Y12Y11Y10Y9 where Y is the complementary nucleotide to X and one or more of Y14, Y13, and Y 12 are optionally part of RTG, and a HSG- TBL in the 5' to 3' direction Zi or Z1Z2 wherein the HSG-TBL is complementary to a portion of the donor site sequence (e.g., XeX? of a donor site sequence) and does not base pair with the target site sequence. In some embodiments, the donor site sequence comprises sequence STIR-Xni-XiX2X3X4X5X6X7X8X9XioXnXi2Xi3Xi4-Xn2-STIR where X is any nucleotide, one or more of X12, X13, and X14 are optionally part of the donor site sequence, STIR is optional, but if present is a sub-terminal inverted repeat comprising 2 to 20 nucleotides, and X8X9 are the core, one or more of X1X2X3X4X5X6X7 are the anti-HSG-donor site sequence, optionally wherein XeX? is the anti-HSG-donor site sequence, and nl and n2 can independently be zero to 10. In some embodiments, the bridgeRNA encodes an LDG in the 5' to 3' directionX1X2X3X4X5X6X7X8 or X1X2X3X4X5X6X7X8X9 and an RDG in the 5' to 3' direction Y14Y13Y12Y11Y10Y9Y8 or Y14Y13Y12Y11Y10Y9 where Y is the complementary nucleotide to X and one or more of Y14, Y13, and Y12 are optionally part of RDG, and a HSG-DBL in the 5' to 3' direction Zi or Z1Z2 wherein the HSG-TBL is complementary to a portion of the target site sequence (e.g., XeX? of a target site sequence) and does not base pair with the target site sequence. In some embodiments, the STIR, if present, comprises a G / T rich nucleotide sequence. In some embodiments, the 5' STIR, if present, comprises a G / T rich nucleotide sequence.

[0022] In some embodiments, the bridgeRNA is engineered, wherein the bridgeRNA comprises one or more engineered HSGs, optionally wherein the bridgeRNA further comprises a wildtype (non-engineered) HSG. In some embodiments, the bridgeRNA has a plurality of HSGs, preferably two HSGs, wherein one or more HSGs of the bridgeRNA with a plurality of HSGs is engineered, and wherein one or more HSGs of the bridgeRNA with a plurality of HSGs is a wildtype (non-engineered) HSG.

[0023] In some embodiments, the nucleotide sequence of the HSG-DBL is engineered to alter its base pairing with said target and / or donor nucleic acid sequences. The engineered HSG-DBL may have increased complementarity with the anti-HSG-target site sequence. The engineered HSG-DBL may have decreased complementarity with the anti-HSG-donor site sequence. In some embodiments, the nucleotide sequence of the HSG-TBL is engineered to alter its base pairing with said target and / or donor nucleic acid sequences. The engineered HSG-TBL may have increased complementarity with the anti-HSG-donor site sequence. The engineered HSG-TBL may have decreased complementarity with the anti-HSG-target site sequence.

[0024] In some embodiments, the one or more engineered HSGs with an engineered nucleotide sequence have an altered interaction between the HSGs and the anti-HSG site sequences. The base pairing between the HSG(s) with an engineered nucleotide sequence and the target and / or donor sequence sites may mediate recombination. The HSG(s) with an engineered nucleotide sequence may possess decreased or no base pairing with between the HSG(s) and the target and / or donor sequence sites before strand exchange. A HSG-DBL with an engineered nucleotide sequence may reduce or avoid pre-strand exchange locking between the HSG-DBL and the anti-HSG-donor site sequence. A HSG-TBL with an engineered nucleotide sequence may reduce or avoid pre-strand exchange locking between the HSG- TBL and the anti-HSG-target site sequence.

[0025] In some embodiments the length of one or more HSG(s) is engineered. The length(s) of the HSG(s) may be reduced. The length(s) of the HSG(s) may be increased. In some embodiments, the HSG-TBL is extended, wherein one or more nucleotides to the 3’ end of the HSG-TBL are modified to base pair with the nucleotides to the 5’ end of the anti-HSG- donor site sequence. In some embodiments, the HSG-DBL is extended, wherein one or more nucleotides to the 3’ end of the HSG-DBL are modified to base pair with the nucleotides to the 5’ end of the anti-HSG-target site sequence.

[0026] In some embodiments, the length(s) of the HSG(s) are increased as a result of substitution, insertion, and or deletion of one or more nucleotides to the 3’ end of the HSG. In some embodiments, the length(s) of the HSG(s) are increased as a result of a combination of two or more of substitution, insertion, and or deletion of one or more nucleotides to the 3’ end of the HSG.

[0027] In some embodiments, the one or more nucleotides to the 3’ end of the HSG are substituted to nucleotides capable of base pairing with the nucleotides to the 5’ end of the post-exchange anti-HSG site sequence. In some embodiments, the one or more nucleotides to the 3’ end of the HSG-DBL are substituted to nucleotides capable of base pairing with the nucleotides to the 5’ end of the anti-HSG-target site sequence. In some embodiments, the one or more nucleotides to the 3’ end of the HSG-TBL are substituted to nucleotides capable of base pairing with the nucleotides to the 5’ end of the anti-HSG-donor site sequence.

[0028] In some embodiments, one or more nucleotides are inserted adjacent to or near the 3’ end of the HSG. In some embodiments, the inserted nucleotides are capable of base pairing with the nucleotides to the 5’ end of the post-exchange anti-HSG site sequence. In some embodiments, the inserted nucleotides reduce binding between the HSG and the nucleotides to the 5’ end of the pre-exchange anti-HSG site sequence. In some embodiments, one or more nucleotides are inserted adjacent to or near the 3’ end of the HSG-DBL, wherein the inserted nucleotides are capable of base pairing with the nucleotides to the 5’ end of the anti-HSG- target site sequence. In some embodiments, one or more nucleotides are inserted adjacent to or near the 3’ end of the HSG-DBL, wherein the inserted nucleotides are capable of reducing base pairing between the nucleotides to the 5’ end of the anti-HSG-donor site sequence. In some embodiments, one or more nucleotides are inserted adjacent to or near the 3’ end of the HSG-TBL, wherein the inserted nucleotides are capable of base pairing with the nucleotides to the 5’ end of the anti-HSG-donor site sequence. In some embodiments, one or more nucleotides are inserted adjacent to or near the 3’ end of the HSG-TBL, wherein the insertednucleotides are capable of reducing base pairing between the nucleotides to the 5’ end of the anti-HSG-target site sequence.

[0029] In some embodiments, one or more nucleotides are deleted adjacent to or near the 3’ end of the HSG. In some embodiments, the nucleotides adjacent to the one or more deleted nucleotides are capable of base pairing with the nucleotides to the 5’ end of the postexchange anti -HSG site sequence. In some embodiments, the nucleotides adjacent to the one or more deleted nucleotides reduce binding between the HSG and the nucleotides to the 5’ end of the pre-exchange anti-HSG site sequence. In some embodiments, one or more nucleotides are deleted adjacent to or near the 3’ end of the HSG-DBL, wherein the nucleotides adjacent to the one or more deleted nucleotides are capable of base pairing with the nucleotides to the 5’ end of the anti-HSG-target site sequence. In some embodiments, one or more nucleotides are deleted adjacent to or near the 3’ end of the HSG-DBL, wherein the nucleotides adjacent to the one or more deleted nucleotides are capable of reducing base pairing between the nucleotides to the 5’ end of the anti-HSG-donor site sequence. In some embodiments, one or more nucleotides are deleted adjacent to or near the 3’ end of the HSG- TBL, wherein the nucleotides adjacent to the one or more deleted nucleotides are capable of base pairing with the nucleotides to the 5’ end of the anti-HSG-donor site sequence. In some embodiments, one or more nucleotides are deleted adjacent to or near the 3’ end of the HSG- TBL, wherein the nucleotides adjacent to the one or more deleted nucleotides are capable of reducing base pairing between the nucleotides to the 5’ end of the anti-HSG-target site sequence.

[0030] In some embodiments, the site sequence(s) is modified. In some embodiments, a target site sequence is engineered, wherein the engineered target site sequence displays improved recombination efficiency with a bridgeRNA comprising one or more wildtype HSGs. In some embodiments, a donor site sequence is engineered, wherein the engineered donor site sequence displays improved recombination efficiency with a bridgeRNA comprising one or more wildtype HSGs.

[0031] In some embodiments, the target site sequence is located on genomic DNA, a linear dsDNA, a dsDNA plasmid, ssDNA or RNA. In some embodiments, the donor site sequence is located on genomic DNA, a linear dsDNA, a dsDNA plasmid, ssDNA or RNA. In some embodiments, the dsDNA plasmid further comprises a polynucleotide sequence for insertion into the donor site sequence. In some embodiments, the dsDNA plasmid further comprises a polynucleotide sequence for insertion into the target site sequence. In some embodiments, thetarget site sequence and donor site sequence on the genomic DNA are located on the same DNA strand. In some embodiments, the target site sequence and donor site sequence on the genomic DNA are located on different chromosomes. In some embodiments, the target site sequence and donor site sequence comprise the same sequence. In some embodiments, the target site sequence and donor site sequence comprise different sequences.

[0032] In some embodiments, the bridgeRNA is a split bridgeRNA. In some embodiments, the bridgeRNA comprises a first RNA molecule that comprises a first portion of the bridgeRNA and a second RNA molecule that comprises a second portion of the bridgeRNA. In some embodiments, the first portion of the bridgeRNA comprises the internal loop comprising the first nucleotide sequence that is complementary to the first target site sequence of the target DNA the second nucleotide sequence that is complementary to a second target site sequence and the handshake guide nucleotide sequence that is complementary to a donor site sequence and the second portion of the bridgeRNA comprises the second internal loop comprising the third nucleotide sequence that is complementary to the first donor site sequence of a donor DNA, and the fourth nucleotide sequence that is complementary to the second donor site sequence, and the handshake guide nucleotide sequence that is complementary to a target site sequence. In some embodiments, the first portion of the bridgeRNA is encoded on a nucleic acid and is operably linked to a first promoter and the second portion of the bridgeRNA is encoded on the same or a different nucleic acid and is operably linked to a second promoter. In some embodiments, the first portion of the bridgeRNA and second portion are encoded on a nucleic acid and further comprise one or more ribozyme sites which results in cleavage of the expressed bridgeRNA into the first and second RNA molecules.

[0033] In some embodiments, any of the LTG, RTG, LDG, and / or RDG of the bridgeRNA are complementary to more nucleotides in their respective target site sequence or donor site sequence than the number of complementary nucleotides in its corresponding naturally occurring bridgeRNA. In some embodiments, the total number of nucleotides comprising the target binding loop and / or donor binding loop of the bridgeRNA is the same as its corresponding naturally occurring bridgeRNA. In some embodiments, the total number of nucleotides comprising the target binding loop and / or donor binding loop of the bridgeRNA is increased as compared to its corresponding naturally occurring bridgeRNA. In some embodiments, any of the LTG, RTG, LDG, and / or RDG of the bridgeRNA are not complementary to a nucleotide of the core sequence in their respective target site sequence ordonor site sequence but are complementary to the same number of nucleotides in the respective target site sequence or donor site sequence as its corresponding naturally occurring bridgeRNA. In some embodiments, the target site sequence comprises sequence X- 1X1X2X3X4X5X6X7X8X9X10X11X12X13X14 where X is any nucleotide, and XSXQ are the core and one or more of X12, X13, and X14 are optionally part of the target site sequence and wherein the bridgeRNA encodes an LTG in the 5' to 3' direction X-1X1X2X3X4X5X6X7 and an RTG in the 5' to 3' direction Y14Y13Y12Y11Y10Y9Y8 where Y is the complementary nucleotide to X and one or more of Y14, Y13, and Y12 are optionally part of RTG, and a HSG-TBL in the 5' to 3' direction Zi or Z1Z2 wherein the HSG-TBL is complementary to a portion of the donor site sequence and does not base pair with the target site sequence. In some embodiments, the donor site sequence comprises sequence STrR-Xni-XiX2X3X4X5X6X7XsX9XioXiiXi2Xi3Xi4- Xn2-STIR where X is any nucleotide, one or more of X12, X13, and X14 are optionally part of the donor site sequence, STIR is optional, but if present is a sub-terminal inverted repeat comprising 2 to 20 nucleotides, and XsX9 are the core, and nl and n2 can independently be 1 to 10, and wherein the bridgeRNA encodes an LDG in the 5' to 3' direction X- 1X1X2X3X4X5X6X7 and an RDG in the 5' to 3' direction Y14Y13Y12Y11Y10Y9Y8 where Y is the complementary nucleotide to X and one or more of Y14, Y13, and Y12 are optionally part of RDG, and a HSG-DBL in the 5' to 3' direction Zi or Z1Z2 wherein the HSG-DBL is complementary to a portion of the target site sequence and does not base pair with the donor site sequence. In some embodiments, one or more of the nucleotides of the LTG, RTG, LDG, RDG, HSG-TBL, and / or HSG-DBL of the bridgeRNA base pair to a nucleotide of their respective target site sequence or donor site sequence via non-canonical base pairing.

[0034] In some embodiments, a scaffold of a bridgeRNA is engineered, wherein the scaffold sequence portion excludes HSGs, RTG, LTG, LDG, and RDG portions. In some embodiments, the bridgeRNA comprising an engineered scaffold is capable of increased bridgeRNA expression, activity, stability, efficiency specificity, and / or imparts novel activities. In some embodiments, the engineered scaffold comprises an engineered ancillary sequence, wherein the engineered ancillary sequence is shortened, extended, truncated, and / or modified. For example, a stem can be shortened from reducing the number of base pairs in the stem (e.g. shortning the stem from 10 base pairs to 5 base pairs). For example, a scaffold can be truncated by removing an entire feature of removing the first or last nucleotides of a feature.

[0035] In some embodiments, the engineered scaffold comprises an engineered TBLstem, wherein the engineered TBL stem is shortened, extended, truncated, and / or modified. In some embodiments, the engineered scaffold comprises an engineered DBL stem, wherein the engineered DBL stem is shortened, extended, truncated, and / or modified. In some embodiments, the engineered scaffold comprises an engineered stem, wherein one or more nucleotides of the engineered stem are modified. In some embodiments, the engineered stem provides enhanced stability of binding to the recombinase, enhanced tetramerization, and / or increased recombination efficiency. In some embodiments, the modification of the engineered stem nucleotides comprises one or more nucleotides inserted into the nonengineered stem sequence, wherein the one or more inserted nucleotides are capable of increasing the total number of base pairs in the stem. In some embodiments, the modification of the engineered stem nucleotides comprises one or more substituted nucleotides, wherein the substituted nucleotides are capable of increasing the strength of base pairing (e.g., by swapping an A:U base pair to a G:C base pair, which has stronger hydrogen bonding). In some embodiments, the engineered scaffold removes a bulge from a non-engineered scaffold. In some embodiments, the engineered scaffold comprises an engineered stem as a result of removing the bulge. In some embodiments, an A:T base pair in a non-engineered stem is replaced with a G:C base pair in the engineered stem.

[0036] In some embodiments, the engineered scaffold comprises an engineered TBL, wherein one or more of the TBL stems, either of the stem flanking the internal loop comprising the programmable nucleotides or the stem of the hairpin flanking the internal loop, is engineered to be shortened, extended, truncated, and / or modified. In some embodiments, the engineered scaffold comprises an engineered DBL, wherein one or more of the DBL stems, either of the stem flanking the internal loop comprising the programmable nucleotides or the stem of the hairpin flanking the internal loop, is engineered to be shortened, extended, truncated, and / or modified. In some embodiments, the engineered scaffold comprises an engineered loop of the DBL stem loop (e.g., the hairpin that is rightmost in a schematic of the DBL - see Figure 3 IB), wherein one or more nucleotides of the loop of the engineered DBL stem loop are modified. In some the engineered loop of the DBL stem loop provides enhanced stability of binding to the recombinase, enhanced tetramerization, and / or increased recombination efficiency.

[0037] In some embodiments, the stem loop of the TBL is replaced with the stem loop sequence of the DBL. In some embodiments, the stem loop of the DBL is replaced with the stem loop sequence of the TBL.

[0038] In some embodiments, the engineered scaffold of the engineered bridgeRNA does not comprise a component present in the non-engineered scaffold. In some embodiments, the engineered bridgeRNA does not comprise a UU bulge, wherein the non-engineered bridgeRNA comprises the UU bulge.

[0039] In some embodiments, the engineered scaffold comprises one or more of the ends of a bridgeRNA, wherein the bridgeRNA is a split-bridgeRNA. In some embodiments one or more of the ends of the bridgeRNA are modified. In some embodiments, the one or more ends of the bridgeRNA are modified chemically or biochemically.

[0040] In some embodiments, the engineered scaffold comprises one or more nucleotides comprising a protecting group. In some embodiments, the protecting group is a pseudoknot. In some embodiments, the protecting group is an mRNA cap. In some embodiments, the protecting group is a tail. In some embodiments, the engineered scaffold comprises one or more DNA nucleotides. In some embodiments, the engineered scaffold comprises one or more non-canonical nucleotides. In some embodiments, the non-canonical nucleotide is inosine. In some embodiments, the engineered scaffold comprises chemical modifications. In some embodiments, the chemical modifications are within the bridgeRNA. In some embodiments, the chemical modifications are at one or more ends of the bridgeRNA. In some embodiments, the chemical modification is a base modification. In some embodiments, the chemical modification is a 2’-3’ cyclic phosphates.

[0041] In some embodiments, the bridgeRNA is a chimeric bridgeRNA. In some embodiments, the chimeric bridgeRNA comprises a first component from a first IS110 ortholog and a second component from a second IS110 ortholog.

[0042] In some embodiments, the bridgeRNA is circular. In some embodiments, a portion of the bridgeRNA is circular. In some embodiments, the TBL of the bridgeRNA is circular. In some embodiments, the DBL of the bridgeRNA is circular. For example, the 5' and 3' end of a bridgeRNA can become closed by connecting the 5' end to the 3' end. In some embodiments, circularization can be applied to an RNA encoding a single binding loop (e.g. TBL or DBL), or a single binding loop (e.g. TBL or DBL) with some extra nucleotides added on the 5' and 3' end so that the circularization point is further away from the loop. In some embodiments, a Twister sister system circularizes the bridgeRNA or the portion of the bridgeRNA. In some embodiments, a hammerhead ribozyme system circularizes the bridgeRNA or the portion of the bridgeRNA.

[0043] In some embodiments, the engineered bridgeRNA recombines a first donor DNAmolecule, a second donor DNA molecule, and a target and donor DNA molecule at a rate different from the rate of recombination by a non-engineered bridgeRNA, wherein the engineered bridgeRNA and the non-engineered bridgeRNA comprise the same HSG(s). In some embodiments, the rate of recombination for the engineered bridgeRNA is greater than the rate of the non-engineered bridgeRNA. In some embodiments, the rate of recombination for the engineered bridgeRNA is lesser than the rate of the non-engineered bridgeRNA.

[0044] In some embodiments, the engineered bridgeRNA comprises a different number of TBLs compared to a non-engineered bridgeRNA, wherein the engineered bridgeRNA and the non-engineered bridgeRNA comprise the same HSG(s). In some embodiments, the engineered bridgeRNA comprises a different number of DBLs compared to a non-engineered bridgeRNA, wherein the engineered bridgeRNA and the non-engineered bridgeRNA comprise the same HSG(s).

[0045] In some embodiments, the engineered bridgeRNA comprises one or more RNA aptamers. In some embodiments, the RNA aptamer is a MS2 stem loop, wherein the MS2 stem loop is capable of recruiting an MCP domain. In some embodiments, the MCP domain is fused to a bridge recombinase.

[0046] In some embodiments, the engineered bridgeRNA has modified tetramerization of the recombination complex. In some embodiments, the engineering bridgeRNA imparts modified properties to the assembly of the recombination complex.

[0047] In certain aspects, the invention provides a recombinant nucleic acid editing system comprising: a) an IS 110 family transposase, or a nucleic acid comprising a sequence encoding the IS110 family transposase; and b) a bridgeRNA, or a nucleic acid comprising a sequence encoding the bridgeRNA, the bridgeRNA comprising at least one stem-loop structure and further comprising at least one internal donor binding loop comprising a first nucleotide sequence that is complementary to a first donor site sequence of a donor DNA, a second nucleotide sequence that is complementary to a second donor site sequence which is on the opposite strand of the donor DNA to the first donor site sequence, and a handshake guide nucleotide sequence that is complementary to a target site sequence, and wherein the bridgeRNA is capable of forming a complex with the IS110 family transposase.

[0048] In certain aspects, the invention provides a recombinant nucleic acid editing system comprising: a) an IS 110 family transposase, or a nucleic acid comprising a sequence encoding the IS110 family transposase; and b) a bridgeRNA, or a nucleic acid comprising a sequence encoding the bridgeRNA, the bridgeRNA comprising at least one stem-loopstructure and further comprising at least one internal donor binding loop comprising a first nucleotide sequence that is complementary to a first donor site sequence of a donor DNA, a second nucleotide sequence that is complementary to a second donor site sequence which is on the opposite strand of the donor DNA to the first donor site sequence, and a handshake guide nucleotide sequence that is complementary to a sequence of a second donor DNA, and wherein the bridgeRNA is capable of forming a complex with the IS110 family transposase.

[0049] In some embodiments, the bridgeRNA does not comprise a target binding loop. In some embodiments, the bridgeRNA comprises a second internal loop comprising a target binding loop, wherein the target binding loop does not bind to any nucleic acids of the nucleic acid editing system. In some embodiments, the IS110 family transposase comprises a RuvC-like DEDD catalytic domain and a transposase domain. In some embodiments, the IS110 family transposase further comprises a linker domain between the RuvC-like DEDD catalytic domain and transposase domain. In some embodiments, the linker domain comprises a coiled-coil linker domain. In some embodiments, the RuvC-like DEDD catalytic domain comprises an amino acid sequence at least 50% identical to a RuvC-like DEDD catalytic domain sequence of SEQ ID NOs: 10176-10523 or 10524-20350. In some embodiments, the IS110 family transposase comprises a RuvC-like DEDD catalytic domain that forms a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621. In some embodiments, the IS110 family transposase RuvC-like DEDD catalytic domain comprises a tertiary structure similar to a tertiary structure of the RuvC-like DEDD catalytic domain of IS621 if the template modeling score (TM-score) for the RuvC-like DEDD catalytic domain of the IS110 family transposase is 0.5 or higher. In some embodiments, the RuvC-like DEDD catalytic domain comprises an amino acid sequence at least 15% identical to a RuvC-like DEDD catalytic domain sequence of SEQ ID NOs: 10176-10523 or 10524- 20350.

[0050] In some embodiments, the transposase domain comprises an amino acid sequence at least 50% identical to a transposase domain sequence of SEQ ID NOs: 10176-10523 or 10524-20350. In some embodiments, the IS110 family transposase comprises a transposase domain that forms a similar tertiary structure to the transposase domain of IS621. In some embodiments, the IS110 family transposase domain comprises a tertiary structure similar to a tertiary structure of the transposase domain of IS621 if the template modeling score (TM- score) for the transposase domain of the IS110 family transposase is 0.5 or higher. In some embodiments, the transposase domain comprises an amino acid sequence at least 15%identical to a transposase domain sequence of SEQ ID NOs: 10176-10523 or 10524-20350. In some embodiments, the IS110 family transposase comprises an amino acid sequence at least 50% identical to SEQ ID NOs: 10176-10523 or 10524-20350.

[0051] In some embodiments, the IS110 family transposase further comprises a tertiary structure similar to a tertiary structure of IS621. In some embodiments, the IS110 family transposase comprises a tertiary structure similar to a tertiary structure of IS621 if the template modeling score (TM-score) for the transposase is 0.5 or higher. In some embodiments, the transposase domain comprises an amino acid sequence at least 15% identical to a transposase domain sequence of SEQ ID NOs: 10176-10523 or 10524-20350. In some embodiments, the IS110 family transposase is an IS 110 group transposase. In some embodiments, the IS110 family transposase is an IS1111 group transposase. In some embodiments, the IS1111 group transposase is IS1111 A or IS1111 229727. In some embodiments, the IS110 group transposase is IS621, ISPal 1, IsPa29, ISMmgl, ISPfl 1 , ISMae40, ISStma6, ISAzs32, ISMex9, ISCARN28, ISAarl6, ISCps7, ISPpu9, ISRel9, ISEsa2, ISMma5, IS900, or ISHne5. In some embodiments, the IS110 group transposase comprises an amino acid sequence at least 50% identical to IS621 (SEQ ID NO: 10176). In some embodiments, the IS110 group transposase comprises an amino acid sequence at least 50% identical to “IS1621 23122” (SEQ ID NO: 10181). In some embodiments, the stem of the stem-loop is 5 to 35 nucleotides long and the loop is 3-10 nucleotides long, and the donor binding loop is 5 to 20 nucleotides long. In some embodiments, the stem of the second stemloop structure comprises 1 to 4 loops or bubbles that are each 1 to 10 nucleotides long.

[0052] In some embodiments, the donor binding loop of the bridgeRNA comprises: a left-donor guide (LDG) comprising, in the 5' to 3' direction, a nucleotide sequence complementary to a first strand of a donor site sequence wherein the 3' end of the LDG complementary to at least one of the nucleotides of a core sequence on the first strand of the donor site sequence; a right-donor guide (RDG) comprising, in the 5' to 3' direction, a nucleotide sequence that is a reverse complement to an opposite strand of the first strand of the donor site sequence wherein the 3' end of the RDG is reverse complementary to at least one of the nucleotides of the core sequence on the opposite strand to the first strand of the donor site sequence; and a handshake guide (HSG-DBL), comprising, in the 5' to 3' direction, a nucleotide sequence that is partially or fully complementary to the opposite strand of the first strand of a target site sequence (anti-HSG-target site sequence) and wherein optionally the HSG-DBL is located 3' to the 3' end of the RDG or wherein the HSG-DBL is located 3'of the 3' end of the RDG and wherein the HSG-DBL is separated from the RDG by 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides; wherein the donor site sequence is a polynucleotide sequence, and the core sequence of the donor site sequence is the same as the core sequence of a target site sequence.

[0053] In some embodiments, the donor binding loop of the bridgeRNA comprises: a left-donor guide (LDG) comprising, in the 5' to 3' direction, a nucleotide sequence complementary to a first strand of a donor site sequence wherein the 3' end of the LDG complementary to at least one of the nucleotides of a core sequence on the first strand of the donor site sequence; a right-donor guide (RDG) comprising, in the 5' to 3' direction, a nucleotide sequence that is a reverse complement to an opposite strand of the first strand of the donor site sequence wherein the 3' end of the RDG is reverse complementary to at least one of the nucleotides of the core sequence on the opposite strand to the first strand of the donor site sequence; and a handshake guide (HSG-DBL), comprising, in the 5' to 3' direction, a nucleotide sequence that is partially or fully complementary to the first strand of a target site sequence (anti-HSG-target site sequence) and wherein optionally the HSG-DBL is located 3' to the 3' end of the RDG or wherein the HSG-DBL is located 3' of the 3' end of the RDG and wherein the HSG-DBL is separated from the RDG by 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides; wherein the donor site sequence is a polynucleotide sequence, and the core sequence of the donor site sequence is the same as the core sequence of a target site sequence.

[0054] In some embodiments, the donor binding loop of the bridgeRNA comprises: a left-donor guide (LDG) comprising, in the 5' to 3' direction, a nucleotide sequence complementary to a first strand of a donor site sequence of a first donor DNA wherein the 3' end of the LDG complementary to at least one of the nucleotides of a core sequence on the first strand of the donor site sequence of the first donor DNA; a right-donor guide (RDG) comprising, in the 5' to 3' direction, a nucleotide sequence that is a reverse complement to an opposite strand of the first strand of the first donor site sequence of the first donor DNA wherein the 3' end of the RDG is reverse complementary to at least one of the nucleotides of the core sequence on the opposite strand to the first strand of the donor site sequence of the first donor DNA; and a handshake guide (HSG-DBL), comprising, in the 5' to 3' direction, a nucleotide sequence that is partially or fully complementary to the opposite strand of the first strand of a donor site sequence of a second donor DNA and wherein optionally the HSG- DBL is located 3' to the 3' end of the RDG or wherein the HSG-DBL is located 3' of the 3'end of the RDG and wherein the HSG-DBL is separated from the RDG by 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides; wherein the donor site sequence of the first donor DNA is a polynucleotide sequence, the donor site sequence of the second donor DNA is a polynucleotide sequence, and the core sequence of the donor site sequence of the first donor DNA is the same as the core sequence of a donor site sequence of the second donor DNA.

[0055] In some embodiments, the donor binding loop of the bridgeRNA comprises: a left-donor guide (LDG) comprising, in the 5' to 3' direction, a nucleotide sequence complementary to a first strand of a donor site sequence of a first donor DNA wherein the 3' end of the LDG complementary to at least one of the nucleotides of a core sequence on the first strand of the donor site sequence of the first donor DNA; a right-donor guide (RDG) comprising, in the 5' to 3' direction, a nucleotide sequence that is a reverse complement to an opposite strand of the first strand of the first donor site sequence of the first donor DNA wherein the 3' end of the RDG is reverse complementary to at least one of the nucleotides of the core sequence on the opposite strand to the first strand of the donor site sequence of the first donor DNA; and a handshake guide (HSG-DBL), comprising, in the 5' to 3' direction, a nucleotide sequence that is partially or fully complementary to the first strand of a donor site sequence of a second donor DNA and wherein optionally the HSG-DBL is located 3' to the 3' end of the RDG or wherein the HSG-DBL is located 3' of the 3' end of the RDG and wherein the HSG-DBL is separated from the RDG by 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides; wherein the donor site sequence of the first donor DNA is a polynucleotide sequence, the donor site sequence of the second donor DNA is a polynucleotide sequence, and the core sequence of the donor site sequence of the first donor DNA is the same as the core sequence of a donor site sequence of the second donor DNA.

[0056] In some embodiments, the HSG-DBL is partially complementary or fully complementary to one or more nucleotides immediately 5' to the core sequence on the opposite strand of the target site sequence or wherein the HSG-DBL is partially complementary or fully complementary to one or more nucleotides located 5' to the core sequence on the opposite strand of the target site sequence and wherein the one or more nucleotides are separated from the core sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides. In some embodiments, the HSG-DBL is partially complementary or fully complementary to one or more nucleotides immediately 5' to the core sequence on the first strand of the target site sequence or wherein the HSG-DBL is partially complementary or fully complementary to one or more nucleotides located 5' to the core sequence on theopposite strand of the target site sequence and wherein the one or more nucleotides are separated from the core sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides. In some embodiments, the HSG-DBL is partially complementary or fully complementary to one or more nucleotides immediately 5' to the core sequence on the opposite strand of the donor site sequence of the second donor DNA or wherein the HSG-DBL is partially complementary or fully complementary to one or more nucleotides located 5' to the core sequence on the opposite strand of the donor site sequence of the second donor DNA and wherein the one or more nucleotides are separated from the core sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides. In some embodiments, the HSG-DBL is partially complementary or fully complementary to one or more nucleotides immediately 5' to the core sequence on the first strand of the donor site sequence of the second donor DNA or wherein the HSG-DBL is partially complementary or fully complementary to one or more nucleotides located 5' to the core sequence on the opposite strand of the donor site sequence of the second donor DNA and wherein the one or more nucleotides are separated from the core sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides.

[0057] In some embodiments, the bridgeRNA is engineered, wherein the bridgeRNA comprises one or more engineered HSGs. In some embodiments, the nucleotide sequence of the HSG-DBL is engineered to alter its base pairing with said target and / or donor nucleic acid sequences. In some embodiments, the target and / or donor site sequence(s) are located on genomic DNA, a linear dsDNA, a dsDNA plasmid, ssDNA or RNA. In some embodiments, the dsDNA plasmid further comprises a polynucleotide sequence for insertion into the donor site sequence of the second donor DNA. In some embodiments, the dsDNA plasmid further comprises a polynucleotide sequence for insertion into the target site sequence. In some embodiments, the target site sequence and donor site sequence on the genomic DNA or the donor site sequence of the first donor DNA and the donor site sequence of the second donor DNA are located on the same DNA strand. In some embodiments, the target site sequence and donor site sequence on the genomic DNA or the donor site sequence of the first donor DNA and the donor site sequence of the second donor DNA are located on different chromosomes. In some embodiments, the target site sequence and donor site sequence comprise the same sequence. In some embodiments, the target site sequence and donor site sequence comprise different sequences. In some embodiments, the donor site sequence of the first donor DNA and the donor site sequence of the second donor DNA comprise the same sequence. In some embodiments, the donor site sequence of the first donor DNA and thedonor site sequence of the second donor DNA comprise different sequences. In some embodiments, the bridgeRNA is a split bridgeRNA.

[0058] In some embodiments, the bridgeRNA comprises a first RNA molecule that comprises the internal donor binding loop comprising a first nucleotide sequence that is complementary to a first donor site sequence of a first donor DNA, a second nucleotide sequence that is complementary to a second donor site sequence which is on the opposite strand of the first donor DNA to the first donor site sequence, and a handshake guide nucleotide sequence that is complementary to a sequence of a second donor DNA; and a second RNA molecule that comprises an internal donor binding loop comprising a first nucleotide sequence that is complementary to a first donor site sequence of the second donor DNA, a second nucleotide sequence that is complementary to a second donor site sequence which is on the opposite strand of the second donor DNA to the first donor site sequence, and a handshake guide nucleotide sequence that is complementary to a sequence of the first donor DNA.

[0059] In some embodiments, the first RNA molecule of the bridgeRNA is encoded on a nucleic acid and is operably linked to a first promoter and the second RNA molecule of the bridgeRNA is encoded on the same or a different nucleic acid and is operably linked to a second promoter. In some embodiments, the first RNA molecule of the bridgeRNA and second RNA molecule are encoded on a nucleic acid and further comprise one or more ribozyme sites which results in cleavage of the expressed bridgeRNA into the first and second RNA molecules. In some embodiments, one or more of the nucleotides of the LDG, RDG, and / or HSG-DBL of the bridgeRNA base pair to a nucleotide of their respective donor site sequences via non-canonical base pairing. In some embodiments, the bridgeRNA comprises a scaffold, wherein the scaffold is engineered, wherein the scaffold sequence portion excludes HSGs, RTG, LTG, LDG, and RDG portions, and optionally wherein the bridgeRNA comprising an engineered scaffold is capable of increased bridgeRNA expression, activity, stability, efficiency specificity, and / or imparts novel activities. In some embodiments, the bridgeRNA is circular or a portion of the bridgeRNA is circular. In some embodiments, the engineered bridgeRNA recombines a first donor DNA molecule and a target DNA molecule and a first donor DNA molecule and a second donor DNA molecule. In some embodiments, the engineered bridgeRNA recombines a first donor DNA molecule and second donor DNA molecule. In some embodiments, the engineered bridgeRNA does not recombine or has reduced recombination of a target DNA molecule and a first donor DNAmolecule.

[0060] In certain aspects, the invention provides a vector comprising any of the nucleic acids of the nucleic acid editing system of the invention. In certain aspects, the invention provides a host cell comprising any of the vector(s) of the invention. In some embodiments, any of the nucleic acids of the nucleic acid editing system further comprise an inducible promoter.

[0061] In certain aspects, the invention provides a method of integrating a DNA molecule of interest into a sequence specific site of a DNA of interest of a cell, the method comprising introducing into the cell: a nucleic acid editing system of the invention.

[0062] In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell.

[0063] In some embodiments, the DNA of interest of the cell comprises a donor site sequence and the DNA molecule of interest is part of a nucleic acid of the nucleic acid editing system comprising a target site sequence. In some embodiments, the DNA of interest of the cell comprises a target site sequence and the DNA molecule of interest is part of a nucleic acid of the nucleic acid editing system comprising a donor site sequence. In some embodiments, the DNA of interest of the cell further comprises a second donor site sequence and the DNA molecule of interest further comprises a second target site sequence and the nucleic acid editing system comprises a second bridgeRNA that targets the second donor site sequence and second target site sequence. In some embodiments, the DNA of interest of the cell comprises a second target site sequence and the DNA molecule of interest further comprises a second donor site sequence and the nucleic acid editing system comprises a second bridgeRNA that targets the second donor site sequence and second target site sequence. In some embodiments, the DNA of interest of the cell comprises a donor site sequence of a second donor DNA and the DNA molecule of interest is part of a nucleic acid of the nucleic acid editing system comprising a donor site sequence of a first donor DNA. In some embodiments, the DNA of interest of the cell comprises a donor site sequence of a donor DNA and the DNA molecule of interest is part of a nucleic acid of the nucleic acid editing system comprising a donor site sequence of a donor DNA. In some embodiments, the sequence of the bridgeRNA was engineered before introduction of the nucleic acid editing system to bind to the donor site sequence and target site sequence. In some embodiments, the DNA of interest of the cell is the genome of the cell. In some embodiments, the DNA of interest of the cell is a plasmid.

[0064] In some embodiments, the donor site sequence or target site sequence of the DNA of interest of the cell has maximal orthogonality to the genome sequence of the cell. In some embodiments, orthogonality is determined using a combination of Hamming distance and Levenshtein distance. In some embodiments, the donor site sequence of the DNA of interest of the cell (if present) comprises an array of donor site sequences comprising repeated identical sequences. In some embodiments, the IS110 family transposase of the nucleic acid editing system is introduced to the cell as a mRNA comprising a sequence encoding the IS110 family transposase; and the bridgeRNA of the nucleic acid editing system is introduced to the cell as a RNA.

[0065] In some embodiments, the target site sequence of the DNA of interest of the cell (if present) comprises an array of target site sequences comprising repeated identical sequences.

[0066] In certain aspects, the invention provides a method of inverting a DNA sequence of a DNA of interest of a cell, the method comprising introducing into the cell: a nucleic acid editing system of the invention, wherein a target site sequence and donor site sequence are present on the same DNA molecule of interest and the LD of the donor site sequence and RT of the target site sequence are on the same DNA strand.

[0067] In some embodiments, the DNA of interest of the cell is the genome of the cell. In some embodiments, the donor site sequence is located closer to the 5’ end of the DNA molecule of interest and the target site sequence is located closer to the 3’ end of the DNA molecule of interest. In some embodiments, the target site sequence is located closer to the 5’ end of the DNA molecule of interest and the donor site sequence is located closer to the 3’ end of the DNA molecule of interest. In some embodiments, the first donor site sequence is located closer to the 5’ end of the DNA molecule of interest and the second donor site sequence is located closer to the 3’ end of the DNA molecule of interest. In some embodiments, the second donor site sequence is located closer to the 5’ end of the DNA molecule of interest and the first donor site sequence is located closer to the 3’ end of the DNA molecule of interest. In some embodiments, the sequence of the bridgeRNA was engineered, before introduction of the nucleic acid editing system, to bind to the donor site sequence and target site sequence. In some embodiments, the target site sequence of the DNA of interest of the cell comprises an array of target site sequences comprising repeated identical sequences. In some embodiments, the donor site sequence of the DNA of interest of the cell comprises an array of donor site sequences comprising repeated identical sequences.

[0068] In certain aspects, the invention provides a method of inverting a DNA sequence of aDNA of interest of a cell, the method comprising introducing into the cell: a nucleic acid editing system of the invention, wherein a first donor site sequence and a second donor site sequence are present on the same DNA molecule of interest and the LD of the first donor site sequence and RD of the second donor site sequence are on the same DNA strand, and the RD of the first donor site sequence and the LD of the second donor site sequence are on the same strand.

[0069] In some embodiments, the DNA of interest of the cell is the genome of the cell. In some embodiments, the first donor site sequence of the DNA of interest of the cell comprises an array of first donor site sequences comprising repeated identical sequences. In some embodiments, the second donor site sequence of the DNA of interest of the cell comprises an array of second donor site sequences comprising repeated identical sequences. In some embodiments, the sequence of the bridgeRNA was engineered, before introduction of the nucleic acid editing system, to bind to the first donor site sequence and second donor site sequence.

[0070] In some embodiments, the IS110 family transposase of the nucleic acid editing system is introduced to the cell as a mRNA comprising a sequence encoding the IS110 family transposase; and the bridgeRNA of the nucleic acid editing system is introduced to the cell as a RNA.

[0071] In certain aspects, the invention provides a method of excising a DNA sequence of a DNA of interest of a cell, the method comprising introducing into the cell: a nucleic acid editing system as described above, wherein a target site sequence and a donor site sequence are present on the same DNA molecule of interest and the LD of the donor site sequence and the LT of the target site sequence are on the same DNA strand. In some embodiments, the DNA of interest of the cell is the genome of the cell. In some embodiments, the sequence of the bridgeRNA was engineered before introduction of the nucleic acid editing system to bind to the donor site sequence and the target site sequence. In some embodiments, the target site sequence of the DNA of interest of the cell comprises an array of target site sequences comprising repeated identical sequences. In some embodiments, the donor site sequence of the DNA of interest of the cell comprises an array of donor site sequences comprising repeated identical sequences.

[0072] In certain aspects, the invention provides a method of excising a DNA sequence of a DNA of interest of a cell, the method comprising introducing into the cell: a nucleic acidediting system described above, wherein a first donor site sequence and a second donor site sequence are present on the same DNA molecule of interest and the LD of the second donor site sequence and the LD of the first donor site sequence are on the same DNA strand.

[0073] In some embodiments, the DNA of interest of the cell is the genome of the cell. In some embodiments, the sequence of the bridgeRNA was engineered before introduction of the nucleic acid editing system to bind to the first donor site sequence and the second donor site sequence. In some embodiments, the first donor site sequence of the DNA of interest of the cell comprises an array of first donor site sequences comprising repeated identical sequences. In some embodiments, the second donor site sequence of the DNA of interest of the cell comprises an array of second donor site sequences comprising repeated identical sequences. In some embodiments, the IS110 family transposase of the nucleic acid editing system is introduced to the cell as a mRNA comprising a sequence encoding the IS110 family transposase; and the bridgeRNA of the nucleic acid editing system is introduced to the cell as a RNA.

[0074] In certain aspects, the invention provides a method of translocating DNA sequences between two linear DNA molecules of interest, the method comprising introducing into a cell: a nucleic acid editing system of the invention, wherein a donor site sequence is present on a first linear DNA molecule and a target site sequence is present on a second linear DNA molecule.

[0075] In some embodiments, the linear DNA molecules of interest of the cell are chromosomes of the cell. In some embodiments, the sequence of the bridgeRNA was engineered, before introduction of the nucleic acid editing system, to bind to the donor site sequence and target site sequence.

[0076] In certain aspects, the invention provides a method of excising a repeat region of a genome of interest comprising: introducing a nucleic acid editing system described above wherein repeat sequences are excised from the genome of interest. In some embodiments, the nucleic acid editing system comprises a bridgeRNA comprising a TBL and a DBL that target the same sequence to mediate recombination reactions between the repeats of the genome of interest. In some embodiments, the nucleic acid editing system comprises a bridgeRNA comprising only a DBL. In some embodiments, the repeat region is of the FXN, HTT, CNBP, or DMPK gene. In some embodiments, the genome of interest is the human genome.

[0077] In some embodiments, the IS110 family transposase of the nucleic acid editing system is introduced as a mRNA comprising a sequence encoding the IS110 family transposase; and the bridgeRNA of the nucleic acid editing system is introduced as a RNA.

[0078] In certain aspects, the invention provides a method of treating a repeat expansion disorder in a subject in need thereof comprising: administering a nucleic acid editing system described above wherein repeat sequences are excised from the genome of the subject. In some embodiments, the nucleic acid editing system comprises a bridgeRNA comprising a TBL and a DBL that target the same sequence to mediate recombination reactions between the repeats of the genome of the subject. In some embodiments, the nucleic acid editing system comprises a bridgeRNA comprising only a DBL.

[0079] In some embodiments, the repeat expansion disorder is Friedreich’s ataxia, Huntington’s disease, or myotonic dystrophy type 1 or type 2. In some embodiments, the subject is a human.

[0080] In some embodiments, the IS110 family transposase of the nucleic acid editing system is administered as a mRNA comprising a sequence encoding the IS110 family transposase; and the bridgeRNA of the nucleic acid editing system is administered as a RNA.

[0081] In certain aspects, the invention provides a method of reverting a pre-existing inversion or translocation of a genome of interest comprising: introducing a nucleic acid editing system described above wherein an inversion or translocation is performed to revert a pre-existing inversion or translocation.

[0082] In some embodiments, the nucleic acid editing system comprises a bridgeRNA comprising a TBL and a DBL that target the breakpoints of the pre-existing inversion or translocation to mediate recombination reactions between the breakpoints of the pre-existing inversion or translocation of the genome of interest.

[0083] In some embodiments, the nucleic acid editing system comprises a bridgeRNA comprising a first DBL and a second DBL that target the breakpoints of the pre-existing inversion or translocation to mediate recombination reactions between the breakpoints of the pre-existing inversion or translocation of the genome of interest.

[0084] In some embodiments, the nucleic acid editing system comprises a bridgeRNA comprising a single DBL that targets the breakpoints of the pre-existing inversion or translocation to mediate recombination reactions between the breakpoints of the inversion or translocation of the pre-existing genome of interest.

[0085] In some embodiments, the genome of interest is the human genome. In some embodiments, the breakpoints have the same sequence. In some embodiments, the breakpoints have different sequences.

[0086] In some embodiments, the IS110 family transposase of the nucleic acid editing system is introduced as a mRNA comprising a sequence encoding the IS110 family transposase; and the bridgeRNA of the nucleic acid editing system is introduced as a RNA.

[0087] In certain aspects, the invention provides a method of treating a pre-existing genomic inversion or translocation in a subject in need thereof comprising: administering a nucleic acid editing system described above wherein an inversion or translocationis performed to revert the pre-existing inversion or translocation of the genome of the subject.

[0088] In some embodiments, the nucleic acid editing system comprises a bridgeRNA comprising a TBL and a DBL that target the breakpoints of the pre-existing inversion or translocation to mediate recombination reactions between the breakpoints of the pre-existing inversion or translocation of the genome of the subject.

[0089] In some embodiments, the nucleic acid editing system comprises a bridgeRNA comprising a first DBL and a second DBL that target the breakpoints of the pre-existing inversion or translocation to mediate recombination reactions between the breakpoints of the pre-existing inversion or translocation of the genome of the subject.

[0090] In some embodiments, the nucleic acid editing system comprises a bridgeRNA comprising a single DBL that targets the breakpoints of the pre-existing inversion or translocation to mediate recombination reactions between the breakpoints of the pre-existing inversion or translocation of the genome of the subject. In some embodiments, the subject is human. In some embodiments, the genome of interest is the human genome. In some embodiments, the breakpoints have the same sequence. In some embodiments, the breakpoints have different sequences.

[0091] In some embodiments, the IS110 family transposase of the nucleic acid editing system is administered as a mRNA comprising a sequence encoding the IS110 family transposase; and the bridgeRNA of the nucleic acid editing system is administered as a RNA.

[0092] In some embodiments, the IS110 family transposase of the nucleic acid editing system in in the form of a mRNA comprising a sequence encoding the IS110 family transposase; and the bridgeRNA of the nucleic acid editing system is in the form of a RNA.

[0093] Other embodiments of the invention are further described in the following sections of the application, including the Drawings, Detailed Description, Examples, and Claims. Still other objects and advantages of the invention will become apparent by those of skill in the art from the disclosure herein, which are simply illustrative and not restrictive. Thus, other embodiments will be recognized by the ordinarily skilled artisan without departing from the spirit and scope of the invention.BRIEF DESCRIPTION OF THE DRAWINGS

[0094] The patent or application file contains at least one drawing originally executed in color. To conform to the requirements for PCT patent applications, many of the figures presented herein are black and white representations of images originally created in color.

[0095] FIGS. 1A-E show general features of IS110 insertion sequence elements. (A) Sequence features of the two groups of IS 110s. The IS110 group is characterized by longer left non-coding ends (LE) and shorter right non-coding ends (RE). The IS1111 group is characterized by shorter LE and longer RE. In both groups, a core sequence motif (l-5nt) is found at both ends of the element. IS 110s were previously thought to lack sub-terminal inverted repeats (STIRs), while IS111 Is were known to have 6-12nt sub-terminal inverted repeats. However, as described herein, most IS 110s also have short STIRs (see Figs. 10A-B). IS110 elements are typically 1000-2000 nt in length. (B) Depiction of domains of IS 110 transposases. IS110 transposases are typically 300-500aa long. The RuvC-like domain (DEDD Tnp ISl 10 by Pfam) includes a canonical DEDD catalytic motif. The IS110 Tnp domain (Transposase_20 by Pfam) has a catalytic serine. (C) Depiction of IS 110 element lifecycle. Genomically integrated IS110 elements cut themselves from the genome and results in scarless repair of the genomic DNA target site and the formation of a circular IS110 element. In the circular form, the RE becomes adjacent to the LE. Insertion can occur into the same dsDNA target site or into new target sites. An inserted linear IS110 element consists of a left non-coding end (LE), coding sequence for a transposase (Tpase), and a right non-coding end (RE). The inserted IS110 element is flanked on the left end with a left flank (LF) (leftmost box) comprising a left target (LT) sequence and on the right end with a right flank (RF) (rightmost box) comprising a right target (RT) sequence. For many IS110 elements, between the LF and LE and between RE and RF are identical “core” sequences (rhombus), although not all IS110 elements may utilize a “core” sequence. IS110 elements excise themselves, resulting in a pre-insertion (“target”) site bearing LF-core (if present)-RF, and a circularelement with RE-core (if present)-LE-Tpase. Concatenation of the RE-LE junction forms a “donor” site sequence as a subsequence of the RE-LE junction, which, if present, includes the other core sequence found on the integrated element. The donor site sequence may also include sub-terminal inverted repeats (STIR) indicated with triangles, although STIRs are not required for IS110 recombinase activity. Concatenation of the RE-LE also forms a promoter which may promote expression of a bridgeRNA from the RE or LE. It may also promote expression of the transposase. The circular form of the element can reinsert into the target site from which it was excised or into any other polynucleotide with a target site sequence; the bridgeRNA encoded within the LE or RE recognizes the donor site sequence and / or the target site sequence to mediate transposition. (D) IS110 transposase phylogenetic tree. IS 110s have several clades but are largely distinguished by the IS110 and IS1111 groups. Showing host kingdom and phylum to demonstrate diverse origins. The location of notable IS110 transposases is highlighted on the tree. (E) Comparison of IS110 group end lengths. IS 110s typically have LEs longer than their REs, while IS111 Is typically have REs longer than their LEs.

[0096] FIGS. 2A-D show identification of the bridgeRNA from the model IS110 IS621. (A) RNAseq of IS110 non-coding ends. A plasmid encoded RE-LE sequence was delivered to E. coli and RNA was extracted and sequenced. Boundaries of an RNA encoded within the LE are defined across 6 orthologs of IS621. (B) Demonstration of bridgeRNA binding to IS621 transposase. The RNA in part A for IS621 was purified and exposed to IS621 transposase at varying concentrations. Microscale thermophoresis is used to measure the binding kinetics of the bridgeRNA to the transposase. A scrambled RNA with no bases matching and a reverse complement of the bridgeRNA serve as negative controls. (C) Determination of IS621 bridgeRNA structure. Hundreds of related LEs of IS621 were aligned and RNA structure was predicted for each. The predominant structure at each position in the alignment was calculated and plotted; structures were characterized as 5' stem, 3' stem, hairpin, other or gap. A structure between the LE start and CDS start emerged that features a consensus bridgeRNA structure and accessory structure on the 5' end. (D) Depiction of IS621 bridgeRNA structure: nnnnnnnnYYnRRnnnnYYYnnnYnnnnRRnnnnYYGGAYGCCGYnYYnRnCCUnnRRYnnn ARYYYGYnnYGUAGAUnnnYGCRnCRRnYRYYnnnnnnnnYnnGYnnnRRRYCGRACnG nAUCnYnGGCYGGYnnnYCGRnARYCYGCAUYACAAGUnGRUnRCRYRAnnnn. The information in C is represented here using the software R2R. The target binding loop anddonor binding loop are labeled. An accessory structure for this particular bridgeRNA is also labeled. The secondary structure of IS621 bridgeRNA structure in Figure 2D is represented in “dot-bracket” notation as:((((((....)))))...) )))..)))).)... where matching parentheses “(“ and “)” indicate base pairs, and unpaired bases are shown as dots

[0097] FIGS. 3A-D show prediction and verification of the mechanism of bridgeRNA recognition of DNA. (A) Depiction of covariation analysis approach. Boundaries of IS 110 elements are identified using comparative genomics. The non-coding ends are concatenated as they would be in the circular form of an IS 110 to identify the donor sites and the preintegration target sites are extracted. bridgeRNA sequences are predicted from non-coding ends. Alignments of target sites (or donor sites) are compared to structurally informed alignments of bridgeRNA sequences to identify covarying positions. (B) Covariation between the bridgeRNA of IS621 and its target and donor. A covariation score as calculated by the software CCMpred was normalized and plotted for each position of the target and donor along the left end. Subsequences of the bridgeRNA, covary with the target and donor. Covarying sequences are observed to be complementary or reverse-complementary to the donor and target. These covarying regions in the bridgeRNA were then inspected for evidence of base-pairing with the target and donor sequences to identify the programmable guide sequences. (C) Model of IS621 bridgeRNA with target and donor sequences. A representation of the R2R structure in Figure 2D is shown with the LTG, RTG, LDG, and RDG positions within the target binding loop and donor binding loop. The LT, RT, LD, and RD are shown with their relative positions to the core sequence found in both the target and donor. Sub-terminal inverted repeats are also shown on the donor. The target binding loop of the bridgeRNA comprises a left-target guide (LTG) and right-target guide (RTG) which are specific for sequences of the target site, i.e., the left target (LT) and right target (RT) sequences, respectively. For IS110 family transposases that use a core sequence, both the LTG and RTG may include base-pairing specificity for at least one base of the core dinucleotide sequence (dashed lines). The donor binding loop of the bridgeRNA comprises a left-donor guide (LDG) and right-donor guide (RDG) which are specific for sequences of the donor site, i.e., the left donor (LD) and right donor (RD) sequences, respectively. For IS110 family transposases that use a core sequence, both the LDG and RDG may include basepairing specificity for at least one base of the core sequence. The donor site may also encodesub-terminal inverted repeats (STIR) which interact directly with the transposase protein, and therefore may be required for the transposition reaction or other parts of the IS110 life-cycle such as cutting and pasting. In some embodiments, additional bridgeRNA nucleotides outside of these described guide sequences can play a role in programming specificity for different target and donor sequences. (D) Demonstration of sequence specific binding of target and donor. IS621 transposase and bridgeRNA ribonucleoproteins were exposed to the WT target and donor sequences as well as scrambled DNA sequences that match at 0 positions. The binding affinity for the WT bridgeRNA for the WT donor and target is shown.

[0098] FIG. 4 shows a diagram of how to reprogram a bridgeRNA to recognize new targets and donors. Depiction of bridgeRNAs programmed to recognize new targets and donors. The WT bridgeRNA is first depicted with the WT target and donor sequence. The target and / or donor loops are modified in each of the following examples. The bridgeRNA sequences are depicted with the LTG, RTG, LDG, and RDG sequences indicated which are capable of basepairing with the target and donor site sequences depicted. These subsequences may be reprogrammed to bind with any desired target or donor sequence. In all five examples, the core nucleotides of the target and donor match each other. In the final example, the STIRs are modified to any sequence, since they are not strictly required for transposition function.

[0099] FIGS. 5A-D shows in cellulo demonstration of transposition and target reprogramming. (A) GFP reporter assay for transposition. A pDonor plasmid encodes an inactive GFP gene adjacent to the RE-LE of IS621, which expresses the bridgeRNA. pDonor is co-transformed into E. coli along with pTarget, which encodes the WT target and the IS621 transposase, both of which are adjacent to promoters. Integration of donor into target activates GFP expression. (B) Demonstration of GFP reporter assay function. E. coli cells are measured for GFP expression on the FITC channel using flow cytometry. GFP expression is observed using the assay in (A) only when the transposase is WT; inactivation of conserved residues in either the RuvC-like domain or Tnp domain abolishes transposition. (C) Diagram of reprogramming bridgeRNAs to recognize new targets. The target loop of the bridgeRNA is modified to recognize a new target. The target is also modified to match the target loop when used in the reporter assay in A. Graph shows demonstration of specific re-targeting of transposition to new targets. Plasmids expressing seven bridgeRNAs with unique target loops were paired with matching targets or the WT target. Transposition is observed only when the matching target is provided, as measured by flow cytometry for GFP expression. (D) Separation of the bridgeRNA from the RE-LE. The bridgeRNA can be separated andexpressed from a separate promoter to achieve higher rates of transposition. The donor sequence can be reduced from 298bp to at least 22bp without affecting transposition efficiency. Truncation to 1 Ibp (removing the STIRs) reduces transposition efficiency to near background. Some integrations are observed when the bridgeRNA is present with the 1 Ibp donor. Systems lacking a bridgeRNA never result in integration.

[0100] FIGS. 6A-E shows IS621 bridgeRNA target / target loop mismatch tolerance and reprogramming screen. (A) Schematic depicting antibiotic resistance reporter design. A minimal donor (22bp) is encoded on a plasmid adjacent to a kanamycin resistance gene. A second plasmid encodes the target, bridgeRNA, and transposase. The target is linked to the bridgeRNA using a barcode. Recombination between the donor and target plasmid results in E. coli survival, and functional bridgeRNA target loop and target pairs are recorded using next generation sequencing. (B) Schematic depicting target specificity screen design. The target and target loop are varied, except for the core of the target and the subsequences of the LTG and RTG that bind the core. The donor loop (and donor) are held constant. Target and target loop pairs are designed to assay single mismatches, double mismatches, and total mismatches. Targets in the screen are selected to reduce the number of off-targets in the E. coli genome. (C) Abundance of target and target loop pairs. Abundance is measured by barcode counts per million reads. Target / target loop pairs with zero mismatches are generally more abundant, while increasing the number of mismatches decreases abundance. (D) Sequence logo of top quintile of targets. The relative enrichment of nucleotides at each position of the target are shown for target / target loop pairs with zero mismatches in the top quintile of 6364 target / target loop pairs. (E) Single mismatch tolerance by position. The relative enrichment of nucleotides for the top quintile of target sets are shown when the target loop does or does not mismatch for each position of the target. The best performing zeromismatch pair in each target set is used to represent the set, and the top quintile of target sets is shown.

[0101] FIGS. 7A-H show IS621 bridgeRNA donor / donor loop mismatch tolerance and reprogramming screen. (A) Schematic depicting antibiotic resistance reporter design. A full length donor is encoded on a plasmid adjacent to a kanamycin resistance gene. The constitutive promoter of the WT system expresses the bridgeRNA. A unique molecular identifier (UMI) identifies a donor / donor loop pair. A second plasmid encodes the target and transposase. Recombination between the donor and target plasmid results in E. coli survival, and functional bridgeRNA donor loop and donor pairs are recorded using next generationsequencing. (B) Schematic depicting donor specificity and reprogramming design. The donor and donor loop are varied, except the core and the subsequences of the LDG and RDG that bind the core. The target loop (and target) are held constant, but the target is a non-WT sequence not found in the E. coli genome. Target and target loop pairs are designed to assay single mismatches and double mismatches for the WT donor sequence. 5000 random perfectly-matched donor and donor loops are also assayed. (C) UMI abundance of donor / donor loop pairs with 1 nt difference from WT. Counts are plotted for pairs with 0 or 1 mismatches between donor and donor loop. UMI abundance of WT donor paired with WT donor loop is depicted as a red dashed line. (D) UMI abundance of donor / donor loop pairs with 2 nt difference from WT. Counts are plotted for pairs with 0 or 2 mismatches between donor and loop. UMI abundance of WT donor paired with WT donor loop is depicted as a red dashed line. (E) UMI abundance of donor / donor loop pairs with 0 mismatches. CPM values of donor / donor loop pairs are binned by the number of nucleotide differences between the reprogrammed donor and WT donor. CPM of WT donor paired with the WT donor loop is depicted as a red dashed line. (F) Single mismatch tolerance by position in the donor. The relative enrichment of nucleotides for the top quintile of donor sets is shown, with all 4x4 = 16 mismatch combinations tested at each position in the donor. (G) Sequence logo of top quintile of donors. The relative enrichment of nucleotides at each position of the target are shown for donor-donor loop pairs with zero mismatches in the top quintile of 5000 donor / donor loop pairs. (H) Demonstration of specific re-targeting of transposition to new donors. Plasmids expressing five bridgeRNAs with unique donor loops were matched with cognate donors or the WT donor. Transposition is observed only when the matching donor is provided, as measured by flow cytometry for GFP expression via FITC. Results were generated using a 22bp donor using the approach in FIG 5D.

[0102] FIGS. 8A-C show a diagram and demonstration of DNA rearrangements with IS621 transposase. (A) Depiction of GFP-reporter assay for DNA insertion. A plasmid encoding a donor and a GFP coding sequence and a plasmid encoding a target plasmid adjacent to a promoter are delivered into E. coli. Co-expression of a bridgeRNA encoding target and donor loops matching the provided target and donor results in efficient insertion in E. coli. (B) Depiction of GFP-reporter assay for excisive recombination of DNA. A plasmid encoding a promoter adjacent to a donor and a target preceded by a terminator and followed by a GFP coding sequence is delivered to E. coli. Co-expression of a bridgeRNA encoding target and donor loops matching the provided target and donor results in efficient excisiverecombination in E. coir, removal of the intervening sequence encoding the terminator enables GFP expression. The reaction results in one DNA molecule becoming two DNA molecules. (C) Depiction of GFP-reporter assay for inversion of DNA. A plasmid encoding a promoter adjacent to a donor and a target preceded by a terminator and GFP coding sequence is delivered to E. coli. Co-expression of a bridgeRNA encoding target and donor loops matching the provided target and donor results in efficient inversion in E. coir, inversion of the sequence between the donor and target enables GFP expression. (A-C) Insertion efficiency is measured by the percent of cells expressing GFP as measured by flow cytometry. Excisive recombination efficiency is measured by the percent of cells expressing GFP as measured by flow cytometry. Inversion efficiency is measured by the percent of cells expressing GFP as measured by flow cytometry.

[0103] FIGS. 9A-B shows a diagram and demonstration of DNA insertion into the E. coli genome with IS621 transposase. (A) Depiction of genome integration assay. An E. coli cell line containing a donor plasmid is made that can grow under kanamycin selection below 37°C due to kanamycin resistance encoded on the donor plasmid. At 37°C, the donor plasmid cannot replicate. A pHelper plasmid encoding IS621 transposase and a bridgeRNA that recognizes the donor on the plasmid and a target in the genome is delivered to E. coli.Growth at 37°C results in selection for A. coli that have integrated the plasmid into the genome, which is required for survival. (B) Integration profile using bridgeRNAs targeting four sites in the genome for integration. Targets are rank ordered by relative abundance of integration locations from nanopore sequencing data. Integration sites are colored by the number of differences between the observed target site and the expected target site.

[0104] FIGS. 10A-B shows identification of sub-terminal inverted repeats in IS110 group IS110 elements. (A) Diagram of approach for identifying sub-terminal inverted repeats. Boundaries of IS 110 elements are identified using comparative genomics and BLAST. The non-coding ends are concatenated as they would be in the circular form and are aligned. Covarying sequences are compared across the donor up to 25bp in each direction from the core. (B) Covariation of sequences within the donor identifies short sub-terminal inverted repeats. A covariation score is plotted for each position of the donor for covariation with itself.

[0105] FIGS. 11 A-C shows prediction and verification of a bridgeRNA expressed from the RE of an IS 1111. (A) Determination of IS 1111 229727 bridgeRNA structure. Hundreds of related REs of IS1111 229727 were aligned and RNA structure was predictedfor each. The predominant structure at each position in the alignment was calculated and graphed; structures were characterized as 5', 3' stem, hairpin, other or gap. A structure between the estimated RE start and element boundary emerges that features a target binding loop and donor binding loop. (B) Depiction of IS1111 229727 bridgeRNA structure nnnnnGn YY n YY GRnRGnGRCGYRGCCCGGY nnnnGnGY A AY CCnCGnnnn YRY nnnGnR RCnRnnn Y nRAYYY nCGnnn Y Y AGAnnGnGGC AGGCnCn Y nnGRCGGAnn Y nnGYRnnG YGGUAnCCARYCCRCGRAURUCAGCnnGAUYnRCCGUCGnnnnnnRCYnGCYnCGYC nCnnYYRRnnRnYnnnnn. The information in A is represented here using the software R2R. The secondary structure of IS1111 229727 bridgeRNA structure in Figure 1 IB is represented in “dot-bracket” notation as:•••))))) ))))) )))•)))))))))))•)))))) )))))••• where matching parentheses “(“ and “)” indicate base pairs, and unpaired bases are shown as dots(C) RNAseq verification of IS1111_229727 bridgeRNA. RNAseq coverage is represented over the RE of IS1111 229727.

[0106] FIGS. 12A-B show alignment of RuvC and Tnp domains of diverse IS110 transposases. (A) Alignment of IS 110 RuvC-like domains. Alignment is depicted with conserved residues and regions. Residues are colored by amino acid chemical properties. (B) Alignment of IS 110 Tnp domains. Alignment is depicted with conserved residues and regions. Residues are colored by amino acid chemical properties.

[0107] FIG. 13 shows diverse predicted bridgeRNA structures associated with diverse IS110 transposases. Showing diverse bridgeRNA consensus structures predicted from across diverse IS110 transposases. The procedure to generate each structure was the same as the procedure used to generate the IS621 bridgeRNA consensus structure. RNA covariance models were clustered using a graph-clustering approach, and consensus structures from 12 different clusters are shown. At least one loop resembling the target and / or donor loop is present in each structure. Significantly co-varying base-pairs are shown with a gray box highlight. The bridgeRNA structures in Figure 13 are representations of sequences with gap positions excluded and trimming of extra unstructured bases.

[0108] FIG. 14A-H shows tertiary structure alignment and analysis of IS 110 transposase proteins. (A) Formula for the Template modeling score (TM-score), where Ltarget is the length of the amino acid sequence of the target protein, and Lcommon is the number of residues that appear in both the template and target structures, di is the distance between theith pair of residues in the template and target structures, and do(Ltarget)=1.24-^(L_target-15)- 1.8 is a distance scale that normalizes distances (Zhang and Skolnick 2004). Alternatively, the score can be normalized according to the length of the query protein, or the score can be normalized by the averaged length of the two proteins. A TM-score has a value in (0,1], and a cutoff of >0.5 is commonly used for identifying proteins with homologous tertiary structures (Zhang and Skolnick 2005). (B) TM-score distribution when aligning predicted IS110 structures to the IS621 AlphaFold structure. Each row shows the distribution of TM-scores when normalized according to the length described on the right - the average of the two lengths, the length of IS621, or the length of the query protein. The dotted line indicates a TM-score of 0.5, a commonly used minimum score threshold for identifying homologous proteins. (C) Structural alignment of two distantly related IS110 proteins. IS621 is shown in green, a separate predicted IS110 transposase structure is shown in cyan. Four different angles of the same structural alignment are shown. These two proteins are 18.1% at the amino acid level, but have a TM-score of 0.805. (D) TM-score distribution of IS 110 structures when clustered and aligned to the IS621 structure. Protein structures were clustered at 100%, 90%, and 50% identity and a representative of each cluster was taken. The TM- score normalized by the average length of the two sequences is shown. Each panel is a different level of percent amino acid identity clustering. (E) TM-scores of RuvC and Tnp domains when aligned to IS621 domains, compared with the full protein TM-scores.Domains were extracted using the boundaries identified by the corresponding Pfam domains (DEDD Tnp ISl 10 and Transposase_20). These domain sub-structures were then aligned to the IS621 sub-structures using TM-align. TM-scores are shown for the full protein, the RuvC domain, and the Tnp domain. (F) IS630 transposase TM-scores vs. IS110 transposase TM- scores. All IS110 family and IS630 family transposase structures were aligned to the IS621 AlphaFold structure. The TM-score normalized by the average length of the two sequences is shown. The IS630 family was selected for comparison because it had a similar protein length distribution to that of IS 110. (G) Schematic demonstrating the location of conserved residues within their respective protein structural domains and the estimated distances between them. On the top panel, showing the 5 conserved residues in the RuvC domain in a representative IS110 structure and a representative IS1111 structure. Residues are colored and labeled with the color red. The 5 positions are labeled P1-P5. Also showing the distances between these residues that are subsequently calculated, including D1-D3, which are colored and labeled as blue, purple, and green, respectively. On the bottom panel, showing the same but for the 5conserved positions in the Tnp domain and the 3 calculated distances. Distances are with respect to the alpha carbon of each residue. (H) Distances between conserved residues in the RuvC and Tnp domains of IS110 AlphaFold structures. Distances were calculated as described in the previous paragraph. Showing here the distribution of distances in angstroms (A) for each distance within each domain. See FIG 14G as a reference for the distances.

[0109] FIG. 15 provides sequence listings for IS110 elements (SEQ ID NOs: 1-348). Elements are represented as 5 '-3' nucleotide sequences in typical FASTA format with additional formatting to indicate subsequences of interest. When available, the annotations include: Dark gray highlighting at the beginning of the sequence indicates the core. The core is only shown once and it is always on the 5' end when annotated. Light gray highlighting indicates the LE and the RE, which always flank the CDS sequence. This is simply defined as the sequence that comes between the CDS and the core or the end of the element. The CDS sequence is shown as non-highlighted sequence with a single underline. The bridgeRNA boundary predictions are shown with lower-case nucleotides. When present, guide sequences are shown with bold typeface. When present, the 4 bold sub-sequences represent the LTG, the RTG, the LDG, and the RDG, in that order. Additional IS110 elements are provided as SEQ ID NOs: 349-10175 of the accompanying sequence listing, which is hereby incorporated by reference in its entirety.

[0110] FIG. 16 provides sequence listings for transposase proteins described herein (SEQ ID NOs: 10176-10523). Proteins are also represented as amino acid sequences in typical FASTA format, with an extra line to represent the secondary structure predictions of each residue. Additional formatting is used to indicate subsequences of interest. When available, the annotations include: Dark gray highlighting to identify the boundaries of the RuvC-like domain as predicted using the DEDD Tnp ISl 10 Pfam domain. Light gray highlighting to identify the boundaries of the Tnp domain as predicted using the Transposase_20 Pfam domain. Bold typeface indicates amino acids that are highly conserved, with up to 5 such amino acids in each domain. The secondary structure prediction was generated using the standard mkdssp tool on all available IS110 transposase AlphaFold structures. These secondary structures were then projected onto sequences in our collection by primary sequence alignment. The different characters indicate: H, Alphahelix; B, Betabridge; E, Strand; G, Helix_3; I, Helix_5; P, Helix PPII; T, Turn; S, Bend; Loop. These secondary structures can be used to orient a person of skill in the art, and be used to identify the coiled-coil linking domain. Additional transposase protein sequences areprovided as SEQ ID NOs: 10524-20350 of the accompanying sequence listing, which is hereby incorporated by reference in its entirety.

[0111] FIG. 17 provides sequence listings for donors (SEQ ID NOs:30354-30529). Donors are represented as 50 nt 5 '-3' nucleotide sequences in typical FASTA format with additional formatting to indicate subsequences of interest. When available, the annotations include: Light gray highlighting indicates the right end (RE) and left end (LE), where the RE is 5' to the core sequence, and the LE is 3' to the core sequence. The core sequence is represented as non-highlighted text with a single underline. When present, the programmable portions of the donor that correspond with the bridgeRNA LDG and RDG are shown with bold typeface. The programmable portion of the donor RE that corresponds with the bridgeRNA LDG is referred to as the left donor (LD) and the programmable portion of the donor LE that corresponds with the bridgeRNA RDG is referred to as the right donor (RD). Additional donor sequences are provided as SEQ ID NOs: 30530-40356 of the accompanying sequence listing, which is hereby incorporated by reference in its entirety.

[0112] FIG. 18 provides sequence listings for targets (SEQ ID NOs: 20351-20526). Targets are represented as 50 nt 5'-3' nucleotide sequences in typical FASTA format with additional formatting to indicate subsequences of interest. When available, the annotations include: Light gray highlighting indicates the left flank (LF) and right flank (RF), where the LF is 5' to the core sequence, and the RF is 3' to the core sequence. The core sequence is represented as non-highlighted text with a single underline. The programmable portions of the target that correspond with the bridgeRNA LTG and RTG are shown with bold typeface. The programmable portion of the target LF that corresponds with the bridgeRNA LTG is referred to as the left target (LT) and the programmable portion of the donor RF that corresponds with the bridgeRNA RTG is referred to as the right target (RT). Additional target sequences are provided as SEQ ID NOs: 20527-30353 of the accompanying sequence listing, which is hereby incorporated by reference in its entirety.

[0113] FIG. 19 provides consensus sequences and structures for bridgeRNA sequences. The name of each model is specified by lines that begin with “>”, just as in a typical FASTA file. The next line is a consensus sequence for the model, where “n” represents any nucleotide, “R” represents an A or G nucleotide, “Y” represents a C or U nucleotide, and then A, C, G, and T represent individual nucleotides. The next four lines indicate 4 possible RNA secondary structures using different confidence thresholds when running the ConsAliFold RNA structure prediction algorithm. These four lines correspond tothe gamma parameters 4, 8, 16, and 32, respectively, with increasing gamma values representing more permissive models (allowing for more structure). The notation used for the secondary structure is referred to as “dot-bracket” notation, where matching parentheses “(“ and “)” indicate base pairs, and unpaired bases are shown as dots (“.”).

[0114] FIG. 20 shows the IS621 transposase AlphaFold model used in the structural analysis. All available IS110 transposase AlphaFold structures were aligned back to this model using the TM-align algorithm to generate TM-scores. This analysis established that a TM-score cutoff of 0.5 is both sensitive and precise for identifying IS110 transposases.

[0115] FIGS. 21A-B show additional examples of predicted bridgeRNA secondary structures with predicted LTG, RTG, LDG, and RDG guide sequences. (A) Showing a schematics of 6 bridgeRNA consensus structures derived from 3 IS 110 group elements and 3 IS1111 group elements. IS110 group elements typically encode their bridgeRNA in the 5' non-coding end (LE) of the element, while IS1111 group elements typically encode their bridgeRNA in the 3' non-coding end (RE). Guide sequences are colored according to the sequence they bind, whether it be the target (blue), the donor (orange), or the core (green). For some members of the IS1111 group, the donor-binding guide sequences are often found within a large multi-loop structure rather than an internal loop. (B) A more detailed representation of the same structures and sequences found in (A). Consensus secondary structures are shown with the IUPAC nucleotide codes circles, colored according to conservation. Highlighted guide sequences are displayed above their corresponding targets and donors for comparison. LTG, RTG, LDG, and RDG are directly labeled. The bridgeRNA structures are representations of sequences with gap positions excluded and trimming of extra unstructured bases.

[0116] FIGS. 22A-E show the utility of extending the natural length of the right target guide (RTG) to increase efficiency and specificity of programmable recombination. (A) Schematic depicting how a longer RTG can be reprogrammed, in addition to how cores can be reprogrammed in conjunction with reprogramming a longer RTG. (B) Relative recombination rate between donors and targets with reprogrammed cores and 4 bp or 7 bp homology RTGs. The assay detailed in FIG 5D was used. Results depict that having longer RTG homology enhances efficiency of recombination with the WT core sequences and reprogrammed core sequences. (C) Schematic depicting approach for genome integration, identical in approach to FIG 9A. (D) On- and off-target integration frequency using 4 base or 7 base RTGs for targeting. The same bridgeRNAs as FIG 9B were utilized to integrate adonor cargo into the E. coli genome with either a 4 base or 7 base RTG. Integration sites were binned by the number of differences from the 11 bp target site sequence intended by the programmed bridgeRNA with a 4 base RTG. (E) Rank order of integration sites averaged over two replicates. The same data as (D) are depicted by the relative number of integrations. High frequency integrations are highlighted by depicting their sequence and to show how the relative targeting specificity is modified when comparing a 4 base and 7 base RTG.

[0117] FIGS. 23A-E show assessment of donor boundaries of an IS 110 bridge recombinase system. (A) Schematic for assaying sequence preference upstream of the donor sequence. The 6 nucleotides upstream of the LD are varied. Recombination is selected for using Kanamycin resistance and successful recombinants are measured via next-generation sequencing (NGS). (B) Schematic for assaying sequence preference downstream of the donor sequence. The 8 nucleotides downstream of the 4th position of the RD are varied, including part of the RD. Assay parameters are otherwise identical to those shown in A. (C) Nucleotide requirement upstream and downstream of the donor sequence. The 5' and 3' STIR sequences are highlighted in pink. (D-E) Sequence preference upstream (D) and downstream (E) of the donor sequence. The 5' and 3' STIR sequences are highlighted in pink.

[0118] FIGS. 24A-C show plasmid-plasmid recombination in human cells. (A) Schematic of plasmid-plasmid recombination assay in human cells. pEffector expressed the bridgeRNA and the recombinase from U6 and Efl a promoters, respectively. pDonor and pTarget are recombined upon co-transfection with pEffector. PCR of the LT-RD junction with primers F and R detect recombination. (B) Verification of plasmid-plasmid recombination. PCR of the LT-RD junction is performed with only pDonor and pTarget, with pEffector lacking a bridgeRNA, and pDonor, pTarget, and pEffector. The recombinase on pEffector was evaluated with three different NLS formats. The recombinase shown is the IS621 recombinase with a bridgeRNA specific for its wild-type donor sequence and a reprogrammed target sequence Target 01. The target binding loop RTG encodes 7bp of homology to the target. (C) Sanger sequencing confirmation of recombination. Sanger sequencing traces are aligned to the entire PCR of the LT-RD junction (top), with a zoomedin version showing the nucleotides proximal to the LT-RD (bottom).

[0119] FIGS. 25A-D shows plasmid inversion in human cells with diverse orthologs. (A) Schematic of plasmid inversion recombination assay in human cells. pEffector expresses the bridgeRNA and the recombinase from U6 and Efl a promoters, respectively. The recombinase is fused to a P2A self cleaving peptide and EGFP to measure recombinaseexpression. PCR of the LT-RD junction with primers F and R detect recombination. (B) Verification of plasmid inversion recombination. PCR of the LT-RD junction is performed in the presence and absence of bridgeRNA for three different NLS configurations for IS621 23122 recombinase. (C) Percentage of cells expressing EGFP 72 hours posttransfection. Four IS110 orthologs are shown each with different NLS configurations. (D) Percentage of mCherry+ cells within the EGFP+ cell population. Four IS110 orthologs are shown each with different NLS configurations. The WT target and donor sequence are recombined for each ortholog. For IS621 127209 and IS621 23122, the sequence flanking the WT 4nt RT was modified to allow 7bp between the RT and the WT RTG.

[0120] FIGS. 26A-D shows bridgeRNA engineering for improved efficiency and specificity. (A) Schematic of the IS110 element IS621 23122 indicating approximate bridgeRNA boundary locations. A bridgeRNA of 179 nt (bRNA179) spans the start of the bridgeRNA to the end of the LE of the element. A bridgeRNA of 260nt (bRNA260) starts at the same location and extends into the CDS of the recombinase. (B) Bridge editing efficiency of an inversion reporter using different length bridgeRNAs. Extending the bridgeRNA to 260nt of natural sequence context increases efficiency relative to the 179nt bridgeRNA. (C) Schematic comparing a WT target binding loop to an LTG-shifted target binding loop. LTG shifting allows targeting of a 16 nt target sequence by binding the 9bp before the core rather than 9 bases including the core, increasing specificity. (D) Bridge editing efficiency of an inversion reporter with a WT bridgeRNA and an LTG-shifted bridgeRNA. Both bridgeRNAs utilize the additional 81nt added to the 3' end of the bridgeRNA in panel b.

[0121] FIGS. 27A-C show engineering of the human genome by delivery of a largeDNA cargo and a bridge recombinase. (A) Schematic depicting bridge editing of the human genome via delivery of a donor plasmid. A recombinase and bridgeRNA specific for the plasmid donor (pDonor, 4.8kb) and the target sequence in the genome results in integration of the donor into the genome. PCR of the LT-RD junction with primers F and R detect recombination. (B) PCR detection of LT-RD junction from genomic DNA. (C) Sanger sequencing confirmation of the integrated donor in the human genome. Sanger sequencing traces are aligned to the entire PCR of the LT-RD junction (top), with a zoomed-in version showing the nucleotides proximal to the LT-RD (bottom).

[0122] FIGS. 28A-E shows engineering of the human genome by delivery of only a bridge recombinase and bridge RNA. (A) Schematic depicting bridge editing of the human genome for inversions via delivery of only recombinase and bridgeRNA. A recombinase andbridgeRNA specific for a genomic donor and genomic target sequence results in inversion when the donor and target are on opposite strands. PCR of the RD-LT junction with primers L and L' and the LD-RT junction with primers R and R' detect recombination. Various orientations of target and donor result in inversion - one is shown here. (B-C) PCR detection of RD-LT and LD-RT for four different bridgeRNAs via agarose gel. The chromosomal locus targeted by the bridgeRNA is shown (above) as well as the relative orientation of the donor and target before and after recombination (bottom). (D) Example of Sanger sequencing confirmation of an inverted locus from panel b. Sanger sequencing traces are aligned to the entire PCR of the RD-LT junction (top left), with a zoomed in version showing the nucleotides proximal to the RD-LT (bottom left). Sanger sequencing traces are aligned to the entire PCR of the LD-RT junction (top right), with a zoomed in version showing the nucleotides proximal to the LD-RT (bottom right). (E) Schematic depicting bridge editing of the human genome for excisions via delivery of only recombinase and bridgeRNA. A recombinase and bridgeRNA specific for a genomic donor and genomic target sequence results in excision when the donor and target are on the same strands. PCR of the LD-RT junction with primers G and G' detect excision from the locus while PCR of the LT-RD junction with E and E' detect the excised DNA. Various orientations of target and donor result in inversion - one is shown here.

[0123] FIGS. 29A-F show engineering of a split bridgeRNA system for recombination. (A) Schematic depicting recombination assay using an LE encoded bridgeRNA specific for a donor and target. (B) Schematic depicting recombination assay using an LE encoded bridgeRNA and a separately expressed target binding loop (TBL). The target binding loop of the LE encoded bridgeRNA has been inactivated by reprogramming the LTG and RTG to have no complementarity to any sequence in the plasmids or organism, while the donor binding loop (DBL) is specific for the donor site sequence. (C) Comparison of recombination efficiency using the structure of the WT bridgeRNA (A) and the split bridgeRNA as depicted in (B). (D) Schematic depicting various split bridgeRNA systems. The bridgeRNA is depicted as a sequence where one or the other binding loop has been reprogrammed to have no specificity for a sequence found in the system. Some versions completely split the bridgeRNA into two, with a separate TBL and DBL. Some versions exclude additional accessory or unstructured nucleotides. The right hand panel indicates what the approximate specificity and structures may be for various iterations of a split bridgeRNA system. (E) Insertion of cargo into the lacZ gene of the E. coli genome. Insertion is performedusing a bridgeRNA with the TBL targeted to the lacZ gene. Blue / white screening of lacZ activity using beta-galactosidase is used to pick colonies bearing the genomic insertion. PCR and agarose gel (right) confirms integration of the cargo into the genome. (F) Insertion of cargo into the lacZ gene of the E. coli genome. Insertion is performed using a bridgeRNA with the TBL reprogrammed to have no specificity for a sequence in the plasmid system or E. coli, while a separate TBL specific for the lacZ gene is expressed from a synthetic promoter such as the system depicted in (B). Blue / white screening of lacZ activity using betagalactosidase is used to pick colonies bearing the genomic insertion. PCR and agarose gel (right) confirms integration of the cargo into the genome.

[0124] FIGS. 30A-B show a summary of mismatch tolerance between an IS 110 bridgeRNA target binding loop and its target. (A) Schematic depicting antibiotic resistance reporter design. A minimal donor (22bp) is encoded on a plasmid adjacent to a kanamycin resistance gene. A second plasmid encodes the target, bridgeRNA, and transposase. The target is linked to the bridgeRNA using a barcode. Recombination between the donor and target plasmid results in E. coli survival, and functional bridgeRNA target loop and target pairs are recorded using next generation sequencing (left). Schematic depicting target specificity screen design. The target and target loop are varied, except for the core of the target and the subsequences of the LTG and RTG that bind the core. The donor loop (and donor) are held constant. Target and target loop pairs are designed to assay single mismatches, double mismatches, and total mismatches. Targets in the screen are selected to reduce the number of off-targets in the E. coli genome. (B) Sequence Abundance of target and target loop pairs. Abundance is measured by barcode counts per million reads. Target / target loop pairs with zero mismatches are generally more abundant, while increasing the number of mismatches decreases abundance. Sequence logo is shown for top quintile of targets. The relative enrichment of nucleotides at each position of the target are shown for target / target loop pairs with zero mismatches in the top quintile of 6364 target / target loop pairs. Mismatch tolerance at each position of the 11 bp target sequence. The x-axis shows the target position, with the CT core held constant. The top panel shows the target nucleotide recovery frequency when the target binding loop contains an A at each guide position as a percentage of recovered recombinants at each position. The second, third, and fourth panel shows the same but when the target-binding loop contains a C, G, or U at each position. Target positions are depicted as the top strand of the DNA.

[0125] FIGS. 31A-B show the bridge recombinase base-pair mechanism. (A) Modelof a bridgeRNA bound to target and donor DNA. A bridgeRNA (bRNA) consists of a 5’ stem loop (5’ SL), a target-binding loop (TBL), and a donor-binding loop (DBL). It binds target DNA (tDNA) and donor DNA (dDNA) via base pairing. The left target guide (LTG) and right target guide (RTG) bind the left target (LT) and right target (RT) of the tDNA, which are on the top strand (TS) and bottom strand (BS) of the dsDNA molecule. The left donor guide (LDG) and right donor guide (RDG) bind the left donor (LD) and right donor (RD) of the dDNA, which are on the TS and DS of the dsDNA molecule. Base pairs within the bRNA and between the bRNA and tDNA / dDNA are based on the IS110 IS621. (B) Diagram of base pairing observed for IS621 in a cryo-EM structure of the recombinase, bRNA, tDNA, and dDNA. The TBL and DBL bound to DNA originate from two different bRNA molecules, rather than one bRNA molecule.

[0126] FIGS. 32A-B show covariation analysis detects potential handshake basepairing mechanism. (A) Nucleotide covariation and base-pairing potential between bRNAs and their target (top) and donor (bottom) sequences, calculated from 2,201 bRNA-target pairs and 5,511 bRNA-donor pairs, as described previously. The IS621 bRNA sequence is shown across the x-axis. Covariation scores calculated from thousands of IS 110 orthologs are colored according to strand complementarity, with -1 (blue) representing high covariation and a bias toward top strand base-pairing, 1 (red) representing high covariation and a bias toward bottom strand base-pairing, and 0 indicating no detectable covariation. Regions with significant covariation signals indicating base-pairing for IS621 are boxed. These data indicated base-pairing potential between TBL-HSG P81 and dDNA P7 (red arrow) and DBL- HSG P166 and tDNA P7 (orange arrow) in IS621, suggesting the functional importance of the base-pairing at these positions across the IS110 orthologs. HSG, handshake guide. (B) Putative schematic of TBL-tDNA and DBL-dDNA before and after the strand exchange.

[0127] FIGS. 33A-D show the effects of handshake base-pairing in vitro and in E. coli. (A) Schematics of the TBL / DBL and tDNA / dDNA sequences used for cryo-EM analysis and in vitro recombination assays. The pre- and post-HSB (handshake base-pairing) bRNAs stabilize the synaptic complex in the pre- and post-strand exchange states, respectively. Mutated nucleotides in the pre- and post-HSB bRNAs and their complementary DNA nucleotides are highlighted. (B) Effects of handshake base-pairing on in vitro DNA activity of IS621. Three bRNAs with different HSG, A81 / U82 / G166 / U167 (WT), G81 / C82 / A166 / U167 (pre-HSB) or A81 / U82 / G166 / C 167 (post-HSB) were used for the in vitro DNA recombination experiments. The tDNA (38 bp) and dDNA (44 bp) substrates werelabeled with FAM and Cy5 at the 5' ends of the top and bottom strands, respectively. The FAM / Cy5-tDNA (38 bp) and FAM / Cy5-dDNA (44 bp) were mixed with the non-labeled dDNA (100 bp) and tDNA (102 bp), respectively. The DNA substrates were incubated with the IS621-bRNA complex at 37°C for 1 h, and the reaction was then analyzed using an 18% TBE-urea gel. Recombination between the FAM / Cy5-dDNA and tDNA and between the FAM / Cy5-tDNA and dDNA yields 60- and 69-bp FAM-labeled products, respectively. The recombination products are marked by asterisks. (C) Schematic depicting locations of reprogrammed handshake bases. (D) Effects of reprogramming handshake base-pairing between the DBL-HSB and dDNA on IS621 -mediated DNA recombination in E. coli using a WT bRNA with a 4nt RTG. Data are shown as mean ± SD for three biological replicates. Schematic depicting assay approach featured in Figure 34A.

[0128] FIGS. 34A-D show the effects of handshake base identity on donor reprogramming. (A) Schematic depicting recombination assay in A. coli. pTarget plasmid encoding the IS621 recombinase and a target is recombined with a pDonor plasmid encoding the donor and bridgeRNA, resulting in GFP expression. (B) Schematic representing the prestrand exchange state of TBL and DBL. The target encoded is a non-WT target, while the donor is the WT donor. Ns within the donor loop feature bases reprogrammed in C and D. The LDG and RDG are always reprogrammed to base pair with the LD and RD, while the HSB is variable. (C) Recombination efficiency with various DBL HSB in combination with reprogrammed donors. Donors with T, G, C, or A in the top strand were tested against the same target with DBL HSB Pl 66 reprogrammed to G, C, U, or A. base pairing between the donor and HSB should inhibit recombination, while base pairing between the HSB and target DNA P7 (C) should enhance recombination. (D) Heat map depicting data from C. Grey boxes indicate a toxic effect on the cells due to base pairing between donor and HSB prior to strand exchange.

[0129] FIGS. 35A-C show the effects of handshake base identity of the target binding loop. (A) Schematic representation of evaluated conditions. Condition 1 assesses recombination between two sequences where P6 and P7 of tDNA and dDNA are the same, and would be base paired before and after recombination by P81 and P82 of the bridgeRNA, respectively. Condition 2 assesses recombination between two sequences where P6 and P7 of tDNA and dDNA are the same, but cannot base pair with P81 and P82 of the bridgeRNA before or after recombination, respectively. In both conditions, there is not handshake between the tDNA and DBL HSGs. (B) List of sequences used in FIG 35C. (C)Recombination efficiency under condition 1 and condition 2. Four pairs of tDNA / dDNA were evaluated. The bridgeRNA LTG / RTG / LDG / RDG base pair with tDNA and dDNA in a typical manner, as depicted in a. P81 and P82 were reprogrammed according to the conditions schematized in FIG. 35 A and detailed in FIG. 35B.

[0130] FIGS. 36A-H show cryo-EM analysis of the IS621 synaptic complex with the pre-HSB bRNA in the pre-strand exchange ‘locked’ state. (A) Single-particle cryo-EM image processing workflow. (B) Euler angle distribution of particles in the final reconstruction. (C) FSC curve. (D) Cryo-EM density map, colored according to the local resolution. (E) Schematic of TBL-tDNA / DBL-dDNA. The base pairs between the HSGs and DNA, which contribute to locking the synaptic complex in the pre-strand exchange state, are highlighted by green lines. (F) Cryo-EM density map for tDNA / dDNA. (G) Structure of TBL- tDNA / DBL-dDNA. (H) Structural comparison between the IS621 synaptic complexes with the WT bRNA and the pre-HSB bRNA in the pre-strand exchange states. DNA cleavage sites are indicated by yellow triangles. Disordered regions in the tDNA and dDNA are depicted as dotted lines and indicated by arrows.

[0131] FIGS. 37A-G show cryo-EM analysis of the IS621 synaptic complex with the post-HSB bRNA in the post-strand exchange Holliday junction state indicating handshake base pairing. (A) Single-particle cryo-EM image processing workflow. (B) Euler angle distribution of particles in the final reconstruction. (C) FSC curve. (D) Cryo-EM density map, colored according to the local resolution. (E) Schematic of TBL-tDNA / DBL-dDNA. The base pairs between the HSGs and DNA, which contribute to locking the synaptic complex in the pre-strand exchange state, are highlighted by green lines. (F) Cryo-EM density map for tDNA / dDNA. (G) Structure of TBL and DBL bound to tDNA / dDNA Holliday junction exhibiting predicted and experimentally demonstrated handshake base pairing.

[0132] FIGS. 38A-C show the effects of handshake base-pairing in human cells. (A) Schematic of plasmid inversion recombination assay in human cells. pEffector expresses the bridgeRNA and the recombinase from U6 and Efl a promoters, respectively. The recombinase is fused to a P2A self cleaving peptide and EGFP to measure recombinase expression. PCR of the LT-RD junction with primers F and R detect recombination. (B) Schematic of the 23122 IS110 bridgeRNA paired with a target and donor sequence, indicating the positions of HSGs in the bridgeRNA and where they bind after strand exchange. (C) Plasmid inversion efficiency with reprogrammed and extended handshake guides.

[0133] FIG. 39 shows the summary of principles for designing handshake guides. A schematic depicts pre- and post- strand exchange base pairing of the TBL and DBL HSGs to tDNA and dDNA. The specific sequences depicted are to demonstrate the base pairing relationships; each quadrant should be considered independently. The descriptions describe the principle that is applied to each HSG / DNA interaction and what should be done for that step of the reaction at the depicted location. In summary, principle 1 describes preventing base pairing between the DBL HSG and the dDNA pre-strand exchange. Principle 2 describes providing base pairing between the TBL HSG and the dDNA post-strand exchange, while optionally avoiding base pairing between the TBL HSG and the tDNA pre-strand exchange. Principle 3 describes providing base pairing between the DBL HSG and tDNA post-strand exchange when possible, which increases recombination efficiency.

[0134] FIGS. 40A-B show mechanism of bridge recombinase mediated programmable DNA recombination. Proposed mechanisms of the IS621 synaptic complex formation (A) and the IS621 -mediated DNA recombination (B). Two IS621 recombinase molecules bind the TBL and DBL in the bRNA, to form the IS621-TBL and IS621-DBL dimeric complexes, respectively. IS621-TBL and IS621-DBL recognize tDNA and dDNA, respectively, and then IS621-TBL-tDNA and IS621-DBL-dDNA form the tetrameric synaptic complex. In the synaptic complex, the top strands of tDNA and dDNA are cleaved at the RuvC / Tnp active sites, with the catalytic S241 residues forming covalent 5'- phosphoserine intermediates. The top strands are then exchanged and re-ligated to form a HJ intermediate, which is then resolved by the cleavage of the bottom strands at the RuvC / Tnp active sites. It is possible that the bottom strands are exchanged, and mismatched nucleotides are excised and repaired in E. coli cells, thereby completing the recombination.

[0135] FIGS. 41A-G show IS621 synaptic complex structure. (A) Schematic of the IS621 insertion sequence element and the domain structure of its cognate recombinase. The CT core dinucleotide sequences are present around the center of the recombination sites of both donor DNA and target DNA. The bRNA encoded in the LE is expressed from the predicted circular form (Durrant, M. G. et al. Bridge RNAs direct modular and programmable recombination of target and donor DNA. Nature in principle accepted (2023)). (B) Schematic of the bRNA bound to the target DNA and donor DNA. 5’ SL, 5’ stem loop; TS, top strand; BS, bottom strand. (C) Nucleotide sequences of the bRNA-complementary regions in the target DNA and donor DNA. Mismatched (MM) nucleotides introduced to the top strands for the structural analysis are shown as lower-case letters. (D, E). Structures of theIS621-bRNA-tDNA-dDNA synaptic complex (D) and the bRNA-tDNA- dDNA complex (E). Disordered regions are indicated by dotted lines, and the CT core dinucleotides (positions 8 and 9) are numbered. In (e), the S241 residues are shown as stick model. (F, G) Structures of the tetramer (F) and monomer (G) of the IS621 recombinase. The catalytic residues are shown as space-filling (F) and stick (G) models. In (G), the core a-helices and p-strands in each domain are numbered. In (D) and (E), DNA cleavage sites are indicated by yellow triangles. In (D) and (F), the active sites with the ordered and disordered S241 residues are indicated by red solid and dashed circles, respectively.

[0136] FIGS. 42A-H show bridgeRNA architecture. (A) Schematic of the bRNA- tDNA-dDNA complex. The covalent 5’-phosphoserine-DNA linkages are indicated by gray lines. Non-canonical base-pairing is indicated by red lines. Disordered nucleotides are indicated by dashed circles. CL, catalytic loop; WED, hydrophobic wedge. (B, C) Structures of TBL-tDNA (B) and DBL-dDNA (C). Disordered regions are indicated by dotted lines. The S241 residues are depicted as space-filling models. In (A-C), DNA cleavage sites are indicated by yellow triangles. (D) Common structural features in tDNA and dDNA. (E-G) Recognition of A67 (E), G72 (F), and the CT core (G). Hydrogen bonds are indicated by green dashed lines. (H) Interaction between the DNA and hydrophobic wedge.

[0137] FIGS. 43A-J show synaptic complex formation. (A, B) Surface representations of the IS621 synaptic complex, colored according to the protomers (A) and domains (B). In (B), the catalytic residues are colored red. (C, D) Active sites formed by RuvC.l / Tnp.4 (C) and RuvC.3 / Tnp.2 (D). The TBL and DBL are shown as space-filling models. DNA cleavage sites are indicated by yellow triangles. Disordered regions are indicated by dotted lines. (E) Locations of the active sites relative to the tDNA and dDNA. (F, G) Close-up views of the active sites formed by RuvC.l / Tnp.4 (I) and RuvC.3 / Tnp.2 (J). Cryo-EM density maps are shown as gray semi-transparent surfaces. The Mg2+ions and water molecules are depicted as cyan and red spheres, respectively. Hydrogen and coordinate bonds are shown as green dashed and solid lines, respectively. DNA cleavage sites are indicated by yellow triangles. (H) Superimposition of the RuvC domains in the four IS621 protomers. The Mg2+ions are depicted as spheres. DNA recombination activities of the WT IS621 and the active-site mutants in the in vitro recombination assays. The tDNA (38 bp) substrate was labeled with Cy5 at the 5’ end of the top strand. The Cy5-tDNA (38 bp) and non-labeled dDNA (100 bp) were incubated with the IS621- bRNA complex at 37°C for 1 h, and then the reaction was analyzed using an 18% TBE-urea gel. Recombination between theCy5-tDNA and dDNA yields a 69-bp Cy5-labeled product. dRuvC,DI 1A / E60A / D102A / D105A. (I) DNA recombination activities of the WT IS621 and the active-site mutants in the bacterial recombination assays. Successful recombination between pTarget and pDonor places the gene encoding green fluorescent protein (GFP) downstream of the synthetic promoter, resulting in fluorescence. IS621, the IS621 recombinase. Data are shown as mean ± SD for three biological replicates.

[0138] FIGS. 44A-I show strand exchange mechanism. (A-I) Schematics of TBL- tDNA / DBL-dDNA, cryo-EM density maps for tDNA / dDNA, and the structures of TBL- tDNA / DBL-dDNA in the IS621 synaptic complexes in the pre-strand exchange state (the complex with the WT bRNA for comparison) (A-C), the post-strand exchange, HJ intermediate state (State 1) (D-F), and the post-strand exchange, HJ resolution state (State 2) (G-I). DNA cleavage sites are indicated by yellow triangles, while partial cleavage and religation are indicated by green triangles.

[0139] FIGS. 45A-B show bridge editing mechanism. (A, B) Proposed mechanisms of the IS621 synaptic complex formation (A) and the IS621 -mediated DNA recombination (B). Two IS621 recombinase molecules bind the TBL and DBL from two different bRNA molecules, to form the IS621-TBL and IS621-DBL dimeric complexes, respectively. IS621- TBL and IS621-DBL recognize tDNA and dDNA, respectively, and then IS621-TBL-tDNA and IS621-DBL-dDNA form the tetrameric synaptic complex. In the synaptic complex, the top strands of tDNA and dDNA are cleaved at the RuvC / Tnp active sites, with the catalytic S241 residues forming covalent 5 ’-phosphoserine intermediates. The top strands are then exchanged and re-ligated to form an HJ intermediate, which is resolved by the cleavage of the bottom strands at the RuvC / Tnp active sites. It is possible that the bottom strands are exchanged, and mismatched nucleotides are excised and repaired in A. coli cells, thereby completing the recombination.

[0140] FIGS. 46A-B show IS621 insertion sequence element. (A) Transposition cycle of the IS621 insertion sequence element. The IS621 elements consist of the left end (LE), the recombinase-coding sequence, and the right end (RE), flanked by the CT core dinucleotide sequences at both ends. The transposition cycle of the IS621 elements consists of excision (generation of a circular form) and insertion (recombination between the circular form and genomic target sites) steps. In the excision step, recombination would occur between the 5’ left target (LT)-core-right donor (RD) region and the 3’ left donor (LD)-core- right target (RT) region in the IS621 locus, resulting in the LT-core-RT region in the originalgenomic site and the LD-core-RD region at the RE-LE junction in the circular form. Importantly, the RE-LE junction contains the reconstituted c70-like promoter sequence, since RE and LE contain a — 35 box and a -10 box, respectively (Durrant, M. G. et al. Bridge RNAs direct modular and programmable recombination of target and donor DNA. Nature in principle accepted (2023)). Thus, the bRNA encoded in the LE is expressed downstream of the core sequence at the RE-LE junction in the circular intermediate (Durrant, M. G. et al. Bridge RNAs direct modular and programmable recombination of target and donor DNA. Nature in principle accepted (2023)). However, it remains unclear whether the excision is mediated by the IS621-bRNA complex and, if so, how the IS621-bRNA complex accomplishes both excision and insertion reactions. For our cryo-EM analysis, the 44-bp donor DNA (the RELE junction with the LD-core-RD sequence in the circular form) and the 38-bp target DNA (the genomic target site with the LT-core-RT sequence) from the natural IS621 element found in E. coh. with the indicated minor modifications were used to assist structural analysis. 5’ SL, 5’ stem loop; TBL, target-binding loop; DBL, donor-binding loop; LTG, left target guide; RTG, right target guide; LDG, left donor guide; RDG, right donor guide; TS, top strand; BS, bottom strand. (B) Schematics showing base-pairing between the bRNA and tDNA / dDNA. Mismatched (MM) nucleotides introduced to the top strands for the structural analysis are shown as lower-case letters.

[0141] FIGS. 47A-K show cryo-EM analysis of the IS621 synaptic complex. (A) Preparation of the IS621-bRNA-dDNA-tDNA synaptic complex. The IS621 recombinase, bRNA, and tDNA / dDNA (containing no mismatch (WT) or 6-nt mismatches (MM) at positions 2-7 in tDNA and dDNA) are mixed, and then purified by a Superose 6 Increase 10 / 300 column. The peak fraction of the synaptic complex with the mismatched DNA (indicated by a red line) was analyzed by SDS- PAGE (10-20%) and TBE-urea PAGE (15%). The proteins and nucleic acids were visualized with CBB and SYBR Gold, respectively. Three slower-migrating bands that may correspond to covalent IS621-DNA intermediates were observed during recombination. (B) In vitro DNA recombination experiments. The 38- bp tDNA and 44-bp dDNA substrates (containing no mismatch (WT) or 6-nt mismatches (MM) at positions 2-7 in their top strands) were incubated with the IS621-bRNA complex at 37°C for 1 h, and then the reaction was analyzed using an 18% TBE-urea gel. The tDNA was labeled with Cy5 at the 5’ end of the top strand. The band intensities of the product and cleaved DNAs were quantified, and the recombination ratios (product DNA / product DNA + cleaved DNA) were calculated. (C) Single-particle cryo-EM image processing workflow. (D)Representative 2D averaged class images. (E) Euler angle distribution of particles in the final reconstruction. (F) Fourier shell correlation (FSC) curves. The map-to-map FSC curve was calculated between the two independently refined half-maps after masking (blue line), and the overall resolution was determined by the gold standard FSC = 0.143 criterion. The map- to-model FSC curve was calculated between the refined atomic models and maps (red line). (G, H) Cryo-EM density maps, colored according to the local resolution (G) and the protein domains (H). In (H), the active sites with the ordered and disordered S241 residues are indicated by red solid and dashed circles, respectively. (I) Effects of 5’ SL deletion on bRNA binding to the IS621 recombinase. Binding of the purified IS621 recombinase to the bRNA, its reverse complement (RC), or the bRNA lacking the 5’ SL (nucleotides A1-U33) (A5’ SL) was analyzed using microscale thermophoresis. Data are shown as mean ± SEM (n = 3). (J) Effects of 5’ SL deletion on IS621 -mediated recombination in A. coli. Data are shown as mean ± SD for three biological replicates. Distance between TBL and DBL. Nucleotides U96-A109 are disordered in the present structure. Modeling of nucleotides U96-A109 (colored gray) suggests that C98 and C99 are -40-A apart, indicating that the TBL and DBL in the synaptic complex structure are derived from two different bRNA molecules.

[0142] FIGS. 48A-D show IS621 recombinase structure. (A) Structural comparison of the RuvC domains of IS621 and Cas9 (PDB: 7S4X). The catalytic residues are shown as stick models. The core a-helices and p-strands are labeled. The RuvC active site of IS621 plays a role in coordinating a Mg2+ion that stabilizes the 5 ’-phosphoserine DNA intermediates, thereby facilitating DNA cleavage and re-ligation. In contrast, most RuvC domains, such as that of Cas9, bind two Mg2+ions and catalyze the DNA cleavage reaction (z.e., the nucleophilic attack of an activated water molecule on the scissile phosphodiester bond in a substrate DNA). (B) Structures of the four IS621 protomers. The disordered S241 loops in IS621.1 and IS621.3 are indicated by dotted lines. (C) Superimposition of the RuvC domains in the four IS621 protomers. (D) Structure of the IS621 tetramer. The active sites formed by RuvC.l / Tnp.4 and RuvC.3 / Tnp.2 are indicated by blue and magenta circles, respectively.

[0143] FIG. 49 shows schematic of interactions between IS621 and nucleic acids. A43 / A67 / A116 / A150 and G48 / G72 / G121 / G155, common structural features in the TBL and DBL, and the core nucleotides (tC8 / tT9 / tA6* / tG7* and dC8 / dT9 / dA6* / dG7*) are highlighted in bold. G48 / G72 / G121 / G155 and U159, which adopt the syn conformation, are shown in italics. The Mg2+ions are depicted as cyan circles. Non-canonical base-pairs are indicated byred squares. The IS621 residues that interact with the nucleic acids through their main chains are shown in parentheses.

[0144] FIGS. 50A-J show DNA recognition mechanism. (A, B) Structures of TBL- tDNA (A) and DBL-dDNA (B). (C, D) Recognition of the CT core sequences in tDNA (C) and dDNA (D). (E, F) Recognition of the CT core-complementary sequences in tDNA (E) and dDNA (F). (G) Modeling of dG8 into the dDNA as part of the core dinucleotide. Interaction between the DNA and the hydrophobic wedge. Similar interactions are observed in the four protomers. In (A-F, and H), cryo-EM density maps are shown as gray semitransparent surfaces. Schematic of the bacterial recombination assays. Successful recombination between pTarget and pDonor places the gene encoding green fluorescent protein (GFP) downstream of the synthetic promoter, resulting in fluorescence. IS621, the IS621 recombinase. (H) DNA recombination activities of the WT IS621 and the hydrophobic-wedge mutants in the bacterial recombination assays. Data are shown as mean ± SD for three biological replicates. dWED, Y264A / M265A / M268A.

[0145] FIGS. 51A-K show synaptic complex formation. (A) Interactions between RuvC.3 and Tnp.2. Similar interactions are observed between RuvC.l and Tnp.4. (b) Interface between the IS621-TBL-tDNA and IS621 -DBL-dDNA dimeric complexes. Nucleotides tT3*-tT5* in the target DNA and nucleotides dA3*-dA6* in the donor DNA are labeled as t3*-t5* and d3*-d6*, respectively. (C) Interactions between RuvC.l and DBL-SL. (D) In vitro DNA recombination activities of IS621 in complex with the WT bRNA or the ADBL-SL bRNA mutant, in which nucleotides C137-G147 were replaced with GAAA. The tDNA (38 bp) substrate was labeled with Cy5 at the 5’ end of the top strand. The Cy5-tDNA (38 bp) and non-labeled dDNA (100 bp) were incubated with the IS621-bRNA complex at 37°C for 1 h, and then the reaction was analyzed using an 18% TBE-urea gel. Recombination between the Cy5-tDNA and dDNA yields a 69-bp Cy5-labeled product. (E) In vitro DNA recombination activities of IS621 in complex with the WT bRNA or separated TBL (nucleotides 31-104) and DBL (nucleotides 99-177). The tDNA (38 bp) substrate was labeled with Cy5 at the 5’ end of the top strand. The Cy5-tDNA (38 bp) and non-labeled dDNA (44 bp) (containing 6-nt mismatches in their top strands) were incubated with the IS621-bRNA complex at 37°C for 1 h, and then the reaction was analyzed using an 18% TBE-urea gel. Recombination between the Cy5- tDNA and dDNA yields a 49-bp Cy5-labeled product. (F) Base-pairing between tDNA and dDNA. In panels A, C, and F, cryo-EM density maps are shown as gray semi-transparent surfaces. (G, H) Structures of the IS621-TBL-tDNA (G) andIS621-DBL-dDNA (H) dimeric complexes. (I) In vitro DNA recombination experiments. The tDNA (38 bp) and dDNA (44 bp) substrates were labeled with Cy5 at the 5’ end of the top strand. The Cy5-tDNA (38 bp) or Cy5-dDNA (44 bp) was mixed with the non-labeled tDNA (102 bp) or dDNA (100 bp), and incubated with the IS621-bRNA complex at 37° C for 1 h. The reaction was then analyzed using an 18% TBE-urea gel. Recombination between the Cy5-dDNA and tDNA and between the Cy5-dDNA and dDNA yields 60- and 64-bp Cy5- labeled products, respectively. In contrast, recombination between the Cy5-tDNA and tDNA and between the Cy5-tDNA and dDNA yields 65- and 69-bp Cy5-labeled products, respectively. For the results with the labeled dDNA, the band intensities of the product and cleaved DNAs were quantified, and the recombination ratios (product DNA / product DNA + cleaved DNA) were calculated. (J, K) Models of the IS621-TBL-tDNA (j) and IS621-DBL- dDNA (k) tetrameric complexes. The models were generated by superimposing IS621-TBL- tDNA and IS621-DBL-dDNA onto IS621- DBL-dDNA and IS621-TBL-tDNA in the synaptic complex.

[0146] FIGS. 52A-B show covariation analysis, (a) Nucleotide covariation and basepairing potential between bRNAs and their target (top) and donor (bottom) sequences, calculated from 2,201 bRNA-target pairs and 5,511 bRNA-donor pairs, as described previously (Durrant, M. G. et al. Bridge RNAs direct modular and programmable recombination of target and donor DNA. Nature in principle accepted (2023)). The IS621 bRNA sequence is shown across the x-axis. Covariation scores calculated from thousands of IS110 orthologs are colored according to strand complementarity, with -1 (blue) representing high covariation and a bias toward top strand base-pairing, 1 (red) representing high covariation and a bias toward bottom strand base-pairing, and 0 indicating no detectable covariation. Regions with significant covariation signals indicating base-pairing for IS621 are boxed. These data indicated base-pairing potential between TBL-HSG P81 and dDNA P7 (red arrow) and DBL-HSG P166 and tDNA P7 (orange arrow) in IS621, suggesting the functional importance of the base-pairing at these positions across the IS110 orthologs. HSG, handshake guide. (B) Schematic of the predicted TBL-tDNA and DBL-dDNA base-pairing patterns before and after the strand exchange.

[0147] FIGS. 53A-C show the effects of handshake base-pairing. (A) Schematics of the TBL / DBL and tDNA / dDNA sequences used for cryo-EM analysis and in vitro recombination assays. The pre- and post-HSB (handshake base-pairing) bRNAs stabilize the synaptic complex in the pre- and post-strand exchange states, respectively. Mutatednucleotides in the pre- and post-HSB bRNAs and their complementary DNA nucleotides are highlighted. (B) Effects of handshake base-pairing on in vitro DNA activity of IS621. Three bRNAs with different HSGs, A81 / U82 / G166 / U167 (WT), G81 / C82 / A166 / U167 (pre-HSB) or A81 / U82 / G166 / C 167 (postHSB), were used for the in vitro DNA recombination experiments. The tDNA (38 bp) and dDNA (44 bp) substrates were labeled with Cy5 and FAM at the 5’ ends of the top and bottom strands, respectively. The Cy5 / FAM-tDNA (38 bp) and Cy5 / FAM-dDNA (44 bp) were mixed with the nonlabeled dDNA (100 bp) and tDNA (102 bp), respectively. The DNA substrates were incubated with the IS621-bRNA complex at 37°C for 1 h, and the reaction was then analyzed using an 18% TBE- urea gel.Recombination between the Cy5 / FAM-dDNA and tDNA and between the Cy5 / FAM-tDNA and dDNA yields 60- and 69-bp Cy5-labeled products, respectively. For the results with the labeled tDNA and pre- and post-HSB bRNAs, the band intensities of the product and cleaved DNAs were quantified, and the recombination ratios (product DNA / product DNA + cleaved DNA) were calculated. (C) Effects of handshake base-pairing between the DBL-HSG and dDNA on IS621 -mediated DNA recombination in E. coli. Data are shown as mean ± SD for three biological replicates.

[0148] FIGS. 54A-H show cryo-EM analysis of the IS621 synaptic complex with the pre-HSB bRNA in the pre-strand exchange ‘locked’ state. (A) Single-particle cryo-EM image processing workflow. (B) Euler angle distribution of particles in the final reconstruction. (C) FSC curve. (D) Cryo-EM density map, colored according to the local resolution. (E) Schematic of TBL-tDNA / DBL-dDNA. The base-pairs between the HSGs and DNA, which contribute to locking the synaptic complex in the pre-strand exchange state, are highlighted by green lines. (F) Cryo-EM density map for tDNA / dDNA. (G) Structure of TBL- tDNA / DBL-dDNA. (H) Structural comparison between the IS621 synaptic complexes with the WT bRNA and the pre- HSB bRNA in the pre-strand exchange states. DNA cleavage sites are indicated by yellow triangles. Disordered regions in the tDNA and dDNA are depicted as dotted lines and indicated by arrows..

[0149] FIGS. 55A-G show cryo-EM analysis of the IS621 synaptic complexes in the post-strand exchange states. (A) Single-particle cryo-EM image processing workflow. (B, E) Euler angle distributions of particles in the final reconstructions of State 1 (B) and State 2 (E). (C, F) FSC curves of State 1 (C) and State 2 (F). (D, G) Cryo-EM density maps of State 1 (D) and State 2 (G), colored according to the local resolution.

[0150] FIGS. 56A-J show structures of the IS621 synaptic complexes in the post-strand exchange states. (A, B) Structures of the IS621 synaptic complexes in the post-strand exchange states (State 1) (A) and (State 2) (B). The active sites with the ordered and disordered S241 residues are indicated by red solid and dashed circles, respectively. (C-J) Close-up views of TS-TBL (State 1) (C), TS-DBL (State 1) (D), TS-TBL (State 2) (E), TS- DBL (State 2) (F), BS-TBL (State 1) (G), BS-DBL (State 1) (H), BS-TBL (State 2) (I), and BS-DBL (State 2) (J). Cryo-EM density maps are shown as gray semi-transparent surfaces. The Mg2+ ions are depicted as cyan spheres. In (C), (D), (I), and (J), the density maps are contoured at two different levels. The top strand is partially re-ligated at the RuvC. l / Tnp.4 active site in State 1 (green arrow) (C), whereas the top strand is fully re-ligated at the RuvC. l / Tnp.4 active site in State 2 (cyan arrow) (E). The top strands are fully re-ligated at the RuvC.3 / Tnp.2 active site in States 1 and 2 (cyan arrows) (D, F). The bottom strands of tDNA and dDNA are not cleaved in State 1 (G, H). The bottom strand of tDNA is fully cleaved at the RuvC.2 / Tnp.3 active site in State 2 (yellow arrow) (I), whereas that of dDNA is partially cleaved at the RuvC.4 / Tnp. l active site in State 2 (green arrow) (J). TS, top strand; BS, bottom strand.

[0151] FIG. 57 shows a comparison of IS621 and Cre. Comparison of the synaptic complex structures, DNA recognition mechanisms, and DNA recombination mechanisms between IS621 and Cre (PDB: 1CRX and 3CRX). In the recombination reactions catalyzed by IS621 and Cre, the top strands in two DNA molecules are cleaved, forming covalent protein-DNA intermediates. DNA cleavage sites are indicated by yellow triangles. Note that the relative angles between the two DNA molecules differ by -180° between the synaptic complexes of IS621 and Cre. TS, top strand; BS, bottom strand.

[0152] FIG. 58 shows regions of bridgeRNA for modification. Schematic of a bridgeRNA highlighting different regions of the scaffold that may be modified to impart modified efficiency and specificity or yield new functions. Schematic is based off of structural relationships within the bridgeRNA, as depicted in more detail in Figures 42A-H and 49.

[0153] FIG. 59 shows structural modifications for increased recombination activity in vitro and in cellulo. In vitro recombination with a split bridgeRNA system, where the WT TBL and various DBLs are provided to donor and target DNA. Lane 1 indicates recombination using a split bridgeRNA with no other modifications to the bridgeRNA scaffold, while lanes 2 and 3 show increased efficiency by increasing the number of base pairs within the DBL stem and by ablating the UU bulge motifs within the DBL loop region,respectively. Lane 4 is a negative control.

[0154] FIGS. 60A-C show engineered bridgeRNA scaffolds for recombination in E. coli. (A) Schematics of engineered bridgeRNAs. The WT bridgeRNA scaffold is depicted with a TBL reprogrammed to a new target (1). Various engineered bridgeRNAs are shown in 2-7 as insets showing the region of modification compared to bridgeRNA 1. (B) Schematic depicting recombination assay in E. coli. pTarget plasmid encoding the IS621 recombinase and a target is recombined with a pDonor plasmid encoding the donor and bridgeRNA, resulting in GFP expression. (C) Recombination efficiency with various engineered bridgeRNAs in E. coli. Recombination efficiency for the bridgeRNAs depicted in (A).

[0155] FIGS. 61A-B show changes to the bRNA scaffold to increase the efficiency of target-donor recombination. (A) Schematics of additional engineered bridgeRNAs evaluated using in vitro. The engineered bridgeRNAs are: 1. 177 nt ncRNA scaffold as described in Durrant & Perry et al., 2023, 2. 258 nt ncRNA where the RNA 3’-end has been extended with additional 81 nt from the genomic context, 3. 177 nt ncRNA modified with the addition of a 50 nt linker between the TBL and DBL loops, 4. 177nt ncRNA modified with the addition of a 100 nt linker between the TBL and DBL loops, 5. 177nt ncRNA modified with the addition of a 150nt linker between the TBL and DBL loops, 6. RNA of 186 nt in length where the DBL stem loop has been appended to the TBL stem loop, 7. RNA of 169 nt in length comprised of two donor-binding loops preceded by the ancillary loop, where the first dbl loop had been recoded to bind the target DNA, 8. As 8 but with an added 50nt linker between the two donor-binding loops. (B) Bar graph showing relative increase in in vitro recombination (IVR) efficiency compared to WT bridgeRNA, assessed using quantitative PCR of the recombination product.

[0156] FIGS. 62A-F show recombination efficiency of a split bridgeRNA in human cells, (a) Schematic of the 23122 WT bridgeRNA. Structure is based on the structure of the IS621 bridgeRNA, which is 87% identical, (b) Schematic of plasmid inversion recombination assay in human cells. pEffector expresses the bridgeRNA and the recombinase from U6 and Efl a promoters, respectively. The recombinase is fused to a P2A self cleaving peptide and EGFP to measure recombinase expression. Inversion activates mCherry expression. Efficiency of recombination is measured as the population positive for mCherry within the population of GFP+ cells, (c) Schematic of a split bridgeRNA for 23122 recombinase. Subsequences of the bridgeRNA constitute a separate target-binding loop (TBL) and donorbinding loop (DBL). (d) Schematic of plasmid inversion recombination assay in human cells.pEffector expresses the either the TBL or the DBL and the recombinase from U6 and Efl a promoters, respectively. Both RNAs are delivered by pooling plasmids encoding either the TBL or the DBL together in transfection. Efficiency of recombination is measured as the population positive for mCherry within the population of GFP+ cells, (e) Recombination efficiency of a single bridgeRNA compared to a split bridgeRNA. Mean fluorescence intensity of mCherry expression (as measured by PE-Texas Red) of the mCherry positive population (mCh+) within the effector positive (Eff+) population (as measured by GFP expression on FITC) is shown. Mean ± SD is shown for 3 biological replicates, (f) Recombination efficiency of a single bridgeRNA compared to a split bridgeRNA. The percentage of mCherry positive cells (as measured by PE-Texas Red) (mCh+) within the effector positive (Eff+) population (as measured by GFP expression on FITC) is shown. Mean ± SD is shown for 3 biological replicates.

[0157] FIGS. 63A-F show relative efficiency of engineered bridgeRNAs in human cells, (a) Schematic of the 23122 WT split target-binding loop. Various engineered versions of this binding loop are paired with the split donor-binding loop shown in (d). (b) Recombination efficiency with various engineered TBLs. All TBLs are paired with the WT split DBL shown in (d). Split bridgeRNA pairs the TBL as shown in (a) and the DBL as shown in (d). The percentage of mCherry positive cells (as measured by PE-Texas Red) (mCh+) within the effector positive (Eff+) population (as measured by GFP expression on FITC) is shown. Mean ± SD is shown for 3 biological replicates, (c) Recombination efficiency with various engineered TBLs. All TBLs are paired with the WT split DBL shown in (d). Split bridgeRNA pairs the TBL as shown in (a) and the DBL as shown in (d). Mean fluorescence intensity of mCherry expression (as measured by PE-Texas Red) of the mCherry positive population (mCh+) within the effector positive (Eff+) population (as measured by GFP expression on FITC) is shown. Mean ± SD is shown for 3 biological replicates, (d) Schematic of the 23122 WT split donor-binding loop. Various engineered versions of this binding loop are paired with the split target-binding loop shown in (d). (e) Recombination efficiency with various engineered DBLs. All DBLs are paired with the WT split TBL shown in (a). Split bridgeRNA pairs the TBL as shown in (a) and the DBL as shown in (d). The percentage of mCherry positive cells (as measured by PE-Texas Red) (mCh+) within the effector positive (Eff+) population (as measured by GFP expression on FITC) is shown. Mean ± SD is shown for 3 biological replicates, (f) Recombination efficiency with various engineered DBLs. All DBLs are paired with the WT split TBL shown in (a). SplitbridgeRNA pairs the TBL as shown in (a) and the DBL as shown in (d). Mean fluorescence intensity of mCherry expression (as measured by PE-Texas Red) of the mCherry positive population (mCh+) within the effector positive (Eff+) population (as measured by GFP expression on FITC) is shown. Mean ± SD is shown for 3 biological replicates.

[0158] FIGS. 64A-E show extension of the TBL Stem for modulating recombination efficiency, (a-c) Schematics of (a) the 23122 WT split target-binding loop, (b) a 23122 split target-binding loop with the addition of bases on the 5' end designed to base pair with the linker region of the bridgeRNA and (c) a 23122 split target-binding loop with a customized stem derived from the linker and donor-binding loop stem, (d) Recombination efficiency with various engineered TBLs with extended stems. All TBLs are paired with the WT split DBL. Split bridgeRNA pairs unmodified TBL and DBL. The percentage of mCherry positive cells (as measured by PE-Texas Red) (mCh+) within the effector positive (Eff+) population (as measured by GFP expression on FITC) is shown. Mean ± SD is shown for 3 biological replicates, (e) Recombination efficiency with various engineered TBLs with extended stems. All TBLs are paired with the WT split DBL. Split bridgeRNA pairs unmodified TBL and DBL. Mean fluorescence intensity of mCherry expression (as measured by PE-Texas Red) of the mCherry positive population (mCh+) within the effector positive (Eff+) population (as measured by GFP expression on FITC) is shown. Mean ± SD is shown for 3 biological replicates.

[0159] FIGS. 65A-E show extension of the DBL Stem for modulating recombination efficiency, (a-c) Schematics of (a) the 23122 WT split donor-binding loop, (b) a 23122 split donor-binding loop with the addition of bases on the 3' end designed to base pair with 5 nucleotides of the linker region of the bridgeRNA and (c) a 23122 split donor-binding loop with the addition of bases on the 3' end designed to base pair with 11 nucleotides of the linker region of the bridgeRNA. (d) Recombination efficiency with various engineered DBLs with extended stems. All DBLs are paired with the WT split TBL. Split bridgeRNA pairs unmodified TBL and DBL. The percentage of mCherry positive cells (as measured by PE- Texas Red) (mCh+) within the effector positive (Eff+) population (as measured by GFP expression on FITC) is shown. Mean ± SD is shown for 3 biological replicates, (e) Recombination efficiency with various engineered DBLs with extended stems. All DBLs are paired with the WT split TBL. Split bridgeRNA pairs unmodified TBL and DBL. Mean fluorescence intensity of mCherry expression (as measured by PE-Texas Red) of the mCherry positive population (mCh+) within the effector positive (Eff+) population (as measured byGFP expression on FITC) is shown. Mean ± SD is shown for 3 biological replicates.

[0160] FIGS. 66A-H show optimization of a split bridgeRNA recombination system, (a) Effects of stacking enhancing mutations to a TBL in a split context and comparison of efficiency to a single bridgeRNA. All TBLs are paired with the WT split DBL. Split bridgeRNA pairs unmodified TBL and DBL. The percentage of mCherry positive cells (as measured by PE-Texas Red) (mCh+) within the effector positive (Eff+) population (as measured by GFP expression on FITC) is shown. Mean ± SD is shown for 3 biological replicates, (b) Effects of stacking enhancing mutations to a TBL in a split context and comparison of efficiency to a single bridgeRNA. All TBLs are paired with the WT split DBL. Split bridgeRNA pairs unmodified TBL and DBL. Mean fluorescence intensity of mCherry expression (as measured by PE-Texas Red) of the mCherry positive population (mCh+) within the effector positive (Eff+) population (as measured by GFP expression on FITC) is shown. Mean ± SD is shown for 3 biological replicates, (c) Effects of stacking enhancing mutations to a DBL in a split context and comparison of efficiency to a single bridgeRNA. All DBLs are paired with the WT split TBL. Split bridgeRNA pairs unmodified TBL and DBL. The percentage of mCherry positive cells (as measured by PE-Texas Red) (mCh+) within the effector positive (Eff+) population (as measured by GFP expression on FITC) is shown. Mean ± SD is shown for 3 biological replicates, (d) Effects of stacking enhancing mutations to a DBL in a split context and comparison of efficiency to a single bridgeRNA. All DBLs are paired with the WT split TBL. Split bridgeRNA pairs unmodified TBL and DBL. Mean fluorescence intensity of mCherry expression (as measured by PE-Texas Red) of the mCherry positive population (mCh+) within the effector positive (Eff+) population (as measured by GFP expression on FITC) is shown. Mean ± SD is shown for 3 biological replicates, (e) Table describing enhanced (enh) binding loops featured in (f) and (g). (f) Comparison of single bridgeRNAs with pairings of optimized split binding loops identifies optimal pairing of enhanced TBL and enhanced DBL. The percentage of mCherry positive cells (as measured by PE-Texas Red) (mCh+) within the effector positive (Eff+) population (as measured by GFP expression on FITC) is shown, (g) Comparison of single bridgeRNAs with pairings of optimized split binding loops identifies optimal pairing of enhanced TBL and enhanced DBL. Mean fluorescence intensity of mCherry expression (as measured by PE- Texas Red) of the mCherry positive population (mCh+) within the effector positive (Eff+) population (as measured by GFP expression on FITC) is shown, (h) Schematic of optimal enhanced TBL (enhTBL-3) and enhanced DBL (enhDBL-2) pairing from panels (f) and (g).

[0161] FIGS. 67A-I show addition of protecting groups to enhanced recombination efficiency, (a-b) Model of a (a) bridgeRNA and (b) bridgeRNA with a structured RNA on the 3' end for the purposes of protecting the bridgeRNA from degradation, (c) Effects of adding various RNA structures to the 3' end of a bridgeRNA. Mean fluorescence intensity of mCherry expression (as measured by PE-Texas Red) of the mCherry positive population (mCh+) within the effector positive (Eff+) population (as measured by GFP expression on FITC) is shown. Mean ± SD is shown for 3 biological replicates, (d) Effects of adding various RNA structures to the 3' end of a bridgeRNA. The percentage of mCherry positive cells (as measured by PE-Texas Red) (mCh+) within the effector positive (Eff+) population (as measured by GFP expression on FITC) is shown. Mean ± SD is shown for 3 biological replicates, (e) List of RNA structures added with sequence listed. Sequences are derived from Nelson et al. 2022 Nat. Biotech, (f-g) Model of a (f) split bridgeRNA and (g) a split bridgeRNA where the 5' and 3' ends are circularized using the Tornado ribozyme system for circularization (Litke et al. 2019 Nat. Biotech.) (h) Effects of circularizing one or both loops of a split bridgeRNA on recombination efficiency. Mean fluorescence intensity of mCherry expression (as measured by PE-Texas Red) of the mCherry positive population (mCh+) within the effector positive (Eff+) population (as measured by GFP expression on FITC) is shown. Mean ± SD is shown for 3 biological replicates, (i) Effects of circularizing one or both loops of a split bridgeRNA on recombination efficiency. The percentage of mCherry positive cells (as measured by PE-Texas Red) (mCh+) within the effector positive (Eff+) population (as measured by GFP expression on FITC) is shown. Mean ± SD is shown for 3 biological replicates.

[0162] FIGS. 68A-C show recombination of identical or unique sequences using bridgeRNAs, split bridgeRNAs, single donor-binding loops, and dual donor-binding loops, (a) Schematic of an inversion reporter assay in human cells. DNA 1 and DNA 2 can be any DNA sequence that a bridgeRNA loop is programmed to bind to. The half-sites indicate the post recombination DNA sequence found surrounding the cores. The effector construct expressing the recombinase may express a bridgeRNA, a single binding loop, or two or more binding loops, (b) Schematics of various bridgeRNA structures, including a split TBL / DBL system and modified DBL structures for performing recombination with only DBL derived RNAs. (c) Recombination efficiency of two identical or unique sequences using various bridgeRNA derived RNAs. (Top) Mean fluorescence intensity of mCherry expression (as measured by PE-Texas Red) of the mCherry positive population (mCh+) within the effectorpositive (Eff+) population (as measured by GFP expression on FITC) is shown. Mean ± SD is shown for 3 biological replicates. (Bottom) The percentage of mCherry positive cells (as measured by PE-Texas Red) (mCh+) within the effector positive (Eff+) population (as measured by GFP expression on FITC) is shown. Mean ± SD is shown for 3 biological replicates. The grey box indicates bridge RNAs with stem loop modifications to prevent DNA2-DNA2 recombination. The graphs show that enhanced DBLs increase donor-donor recombination.

[0163] FIGS. 69A-D show insertion of a donor plasmid into the human genome using donor-binding loop mediated recombination, (a) Schematic of an insertion assay in human cells. A construct expresses the 23122 recombinase and the WT bridgeRNA. A donor plasmid encodes the WT donor sequence and an expression cassette for puromycin resistance tied to mCherry expression with a P2A reporter. The two plasmids are transfected into a human cell bearing many locations in the genome identical to the WT donor sequence. After selection for 14 days with puromycin, cells with insertions are isolated, genomic DNA extracted, and insertion sites mapped using a previously reported modified UDiTaS method (Durrant, Fanton, Tycko et al. 2022 Nat. Biotech.), (b) Sequence logo of all genomic insertions. Sequence logo is representative of 3 biological replicates, (c) General types of donor-like genomic sequences observed to have insertions mediated by donor-binding loop insertion. The 14mer and 1 Imer match the WT donor binding loop sequences, and n is the number of times these sequences appear in the reference genome. Shifted sequences are similar to 1 Imer, except bearing one or more insertions immediately upstream of the core. Underlines indicate nucleotides likely base-paired with by the LDG. H is IUPAC code indicating A or C or T. (d) Proportion of insertions by sequence type. All insertions across three biological replicates are summarized. Only exact matches to the sequences in (c) are shown.

[0164] FIGS. 70 shows in vitro recombination using a recombinase and bridgeRNA from different orthologs. Measurement of in vitro recombination (IVR) efficiency using qPCR. The recombinase and bridgeRNA used in the experiment is indicated; the same target and donor is used for all assays.

[0165] FIG. 71 shows IS110 family recombinases. Multiple sequence alignment of the IS110 family recombinases. Key residues are indicated by triangles. The figure was prepared using Clustal Omega (www.ebi.ac.uk / Tools / msa / clustalo) and ESPript3 (espript.ibcp.fr / ESPript / ESPript).

[0166] FIG. 72 shows a comparison of IS621 with other recombinases and transposases. Comparison of the primary and tertiary structures and the DNA recombination mechanisms among the DEDD recombinase IS621, the tyrosine recombinase Cre (PDB: 1CRX), the serine recombinase Bxbl, and the DDE transposase IstA (PDB: 8B4H). The crystal structure of the resolvase (PDB: 1ZR2) is shown as the representative structure of a serine recombinase, since no structural information is available for Bxbl. For comparison, the two DNA molecules bound to Cre and the y5 resolvase are labeled as tDNA and dDNA and colored blue and magenta, respectively, although they catalyze recombination between two DNA molecules with the identical recognition sequences (ZoxP for Cre and res for the y5 resolvase). The IstA structure binds a donor DNA (the inverted terminal repeat of the IS21 element), but lacks a target DNA. The four protomers in the synaptic complexes are numbered, and the active sites are marked by yellow circles. NTD, N-terminal domain; CTD, C- terminal domain; RD, recombinase domain; ZD, zinc ribbon domain; HTH, helix-tum- helix.

[0167] FIGS. 73A-L shows Human Genome Insertion with Enhanced Bridge RNAs. (a) Detailed diagram of 23122 WT bRNA and representative schematic, (b) Detailed diagram of 23122 WT Split bRNA and representative schematic (c) Detailed diagram of 23122 enhanced split bRNA and representative schematic (d) Detailed diagram of 23122 enhanced single bRNA and representative schematic (e) Schematic of inversion assay for measuring efficiency of enhanced bridge RNAs. Inversion results in mCherry expression, (f) Inversion efficiency of WT and enhanced bridge RNA scaffolds. The 5' stem loop may be added to split bridge RNA formats to enhance efficiency, and the modifications engineered as a split bridge RNA also enhances recombination activity when utilized as a single bridge RNA. (g) Diagram for genome insertion. A plasmid bearing a human genome orthogonal (HGO) DNA sequence is delivered to human cells for insertion into a human genome sequence (HGS) that appears once in the human genome. Insertion efficiency is measured by performing digital droplet PCR (ddPCR) on the insertion junction, (h) Schematic of possible bRNA scaffolds for recombination of HGO and HGS, where the TBL binds HGS and the DBL binds HGO. (i) Schematic of possible bRNA scaffolds for recombination of HGO and HGS, where the DBL binds HGS and the TBL binds HGO. (j) Insertion efficiency into HGS7 with various bRNA scaffolds, (k) Insertion efficiency into HGS26 with various bRNA scaffolds. (1) Insertion efficiency using only DBLs, with two DBLs delivered to cells, with one specific for HGS, and one specific for HGO

[0168] FIGS. 74A-E show high throughput insertion into human genes, (a) Schematic detailing human genome site selection strategy and molecular readout strategy. Human genes from mostly non-essential genes were chosen based on Funk et al. Cell 2022. These genes were filtered for expression level using transcript per million (TPM) information from Katrekar et al. ELife 2022. Unique 14mer sequences, which appear only once in the human genome, were selected within non-essential genes, provided they appeared within 5kb of the exon 1 start site. Various human genome orthogonal sequences (HGO) were paired with these human genome sequences (HGS) based on core compatibility, providing identical core sequences between the HGO and HGS. Bridge RNAs were designed such that the TBL recognizes the plasmid bearing HGO and the DBL recognizes the HGS in the genome. Amplification of the insertion junction followed by next-generation sequencing results in a count of unique molecular identifiers (UMI), with each UMI corresponding to a unique insertion event, (b) Replicate correlation for UMIs across all different genes. Analysis required alignment of at least 4000 reads to the reference junction sequence, consisting of the half-site between the intended human genome site and the plasmid, (c) Insertion efficiency by gene, based on UMIs per 1000 reads, (d) Schematic depicting two options for genome insertion, one where the DBL binds the genome and the TBL binds the plasmid, and the other where the DBL binds the plasmid and the DBL binds the genome, (e) Insertion efficiency into select genes as measured by digital droplet PCR (ddPCR).

[0169] FIGS. 75A-C show specificity analysis of genome insertion with 23122. (a) Schematic of specificity analysis strategy. A plasmid bearing a human genome orthogonal sequence and a UMI is inserted into the genome, with possibilities for on- or off-target insertion. Introduction of a constitutive puromycin resistance cassette enables positive selection for cells bearing genomic insertions over a period of 18 days. After enrichment, tagmentation is used to unbiasedly amplify all on- and off-target insertions for next generation sequencing, according to the protocol for Uditas featured in Durrant, M.G., Fanton, A., Tycko, J. et al., Systematic discovery of recombinases for efficient integration of large DNA sequences into the human genome. Nat Biotechnol 41, 488-499 (2023), doi.org / 10.1038 / s41587-022-01494-w, the contents of which is hereby incorporated by reference in its entirety, (b) Insertion specificity across seven genomic sites. Each panel is separated by insertions mediated by an enhanced single bridge RNA where the DBL binds the genome and the TBL binds the plasmid, or the DBL binds the plasmid and the TBL binds the genome. The percentage of UMIs found at the on-target site is shown in green, while thepercentage of UMIs found at off-target sites is shown in gray. Each dot represents a human genome site where insertion was found, (c) Schematic detailing the off-target landscape for genome insertion. Recombination between an HGO and HGS or HGO and an HGS-like sequence requires both a TBL and DBL to occur. For recombination between HGO and an HGO-like sequence, only a DBL is required, since the DBL is able to mediate recombination between DNA in the absence of a TBL. Off-target rearrangements within the genome may also occur when performing genome insertion, in addition to multimerization of the plasmid being delivered. Importantly, this off-target landscape, in addition with specificity data, indicates that it is optimal to have the DBL programmed to bind the genome, and the TBL programmed to bind the plasmid, in order to limit recombination reactions that can occur with only a DBL.

[0170] FIGS. 76A-D show human genome rearrangements with enhanced bridge RNAs. (a) Schematic of performing excision within the human genome. An enhanced bridge RNA may be designed in at least two ways, with one of the TBL or DBL programmed to bind an “anchor” sequence and one of the TBL or DBL programmed to bind a “partner” sequence. Recombination between anchor and partner results in the excision and effective deletion of the intervening sequence, presumably also yielding an extrachromosomal circular DNA fragment, (b) Schematic of performing inversion within the human genome. An enhanced bridge RNA may be designed in at least two ways, with one of the TBL or DBL programmed to bind an “anchor” sequence and one of the TBL or DBL programmed to bind a “partner” sequence. Recombination between anchor and partner results in the inversion of the intervening sequence, (c) Schematic of molecular readout for measuring excision and inversion. Tagmentation with Tn5 transposase unbiasedly tags genomes that do or do not bear a rearrangement. Unbiased amplification between a primer flanking the rearrangement and a primer binding the tagmented DNA yields DNA fragments that can be sequenced with nextgeneration sequencing. The ratio of modified junctions to all sequenced junctions is the percentage of genomes bearing a rearrangement, (d) Results of various excision and inversion reactions within the human genome. The distance between anchor and partner sequence is shown, along with the bridge RNA configuration used for the rearrangement. Importantly, the configuration of the bridge RNA determines the efficiency of the recombination between the two sites, with one configuration, determined empirically, having superior rearrangement efficacy over the other.

[0171] FIGS. 77A-L show expanded repeat compression with Enhanced BridgeRNAs and Donor Binding Loops, (a) Schematic of the FXN gene, with region of repeat expansion shown. Patients with the disease Friedreich’s ataxia show repeat expansion from the WT sequence, (b) Schematic of recombination within a repeat sequence of GAA. A target and donor can be specified within the repeat region, such that a suitable NT core sequence is recognized on either strand (in the case of FXN, CT). Recombination between target and donor results in the excision of the expanded repeat, which is restorative of a nondisease phenotype, (c) Schematic of using an inversion assay to assess recombination between identical sequences, (d) Schematic of GAA repeat sequences separated from one another on the reporter plasmid. In one assay, five GAA repeats are provided, while in the second 11 GAA repeats are provided. Each additional repeat yields an additional register for which a bridge RNA to bind, (e) Schematic of a bridge RNA programmed to recognize a GAA repeat sequence. Both TBL and DBL have specificity for the same 14nt sequence within a GAA repeat sequence, base-pairing with the reverse complement of a subsequence of GAA repeat TCTTCTTCTTCTTC. (f) Recombination efficiency between GAA repeat sequences. With a single bRNA, weak recombination is observed, although more recombination is observed when the sequence has multiple registers via which the bridge RNA can recognize DNA. Using a split bridge RNA featuring only a DBL that recognizes the GAA repeat sequence results in much more robust recombination, (g) Schematic of recombination within a repeat sequence CAG, or its reverse complement, CTG. These repeats are found in the repeat expansion disorders Huntington’s disease and myotonic dystrophy type 1. Recombination between target and donor results in the excision of the expanded repeat, which is restorative of a non-disease phenotype, (h) Recombination efficiency between CAG / CTG repeat sequences. With a single bRNA, weak recombination is observed, although more recombination is observed when the sequence has multiple registers via which the bridge RNA can recognize DNA. Using a split bridge RNA featuring only a DBL that recognizes the GAA repeat sequence results in much more robust recombination, (i) Schematic of recombination within a repeat sequence CCTG. These repeats are found in the repeat expansion disorder myotonic dystrophy type 2. Recombination between target and donor results in the excision of the expanded repeat, which is restorative of a non-disease phenotype, (j) Recombination efficiency between CCTG repeat sequences. With a single bRNA, weak recombination is observed, although more recombination is observed when the sequence has multiple registers via which the bridge RNA can recognize DNA. Using a split bridge RNA featuring only a DBL that recognizes the GAA repeat sequence results in muchmore robust recombination, (k) Change in absolute number of repeats observed at recombined junctions using premium PCR next generation sequencing from Plasmidsaurus. The inverted junction of the recombined plasmid in c) was sequenced and the number of repeats observed at the junction were counted and compared to the number of repeats at the unrecombined junction (11 or 12). At least 3000 reads were considered in the analysis. (1) Percentage of total number of repeats expected to be excised in an excision reaction. The change in number of repeats was calculated as a percentage of all reads. Two instances of 11 or 12 repeats were recombined with one another at a distance. In the case they were on the same strand, reduction of two 11 repeat arrays to a single array of 11 repeats represents 50% of the repeats being mobilized.DETAILED DESCRIPTION

[0172] The present invention relates to the IS110 transposon family. The IS110 transposons encode both a “bridgeRNA” molecule and a transposase protein. The bridgeRNA molecule in concert with the transposase mediates site-specific recombination between one or more DNA molecules containing a target site sequence and a donor site sequence. The target site sequence and the donor site sequences can be on the same DNA molecule or different DNA molecules. Generally, as used herein, the target site and donor site sequences are simply nucleic acid sequences that associate with, or are recognized by, an IS110 bridgeRNA and transposase complex, and depending on the orientation of these sequences and whether these sequences are on the same or different molecules, a transposition reaction will occur resulting in recombination between the target site and donor site sequences such that the result is either an insertion (or translocation), excisive recombination, or inversion. Thus, for insertion or translocation reactions, the target sequence and the donor sequence are on different molecules. For excisive recombination and inversion, the target and donor site sequences are on the same molecule, and depending on the orientation of the target and donor site sequences, intervening sequences are excised or inverted. Such recombination reactions may be employed to recombine any DNA sequence with any other DNA sequence in a programmable manner, without any requirements to use DNA sequences originating from the IS110 element. More specifically, the present invention provides recombinant IS110 transposons where the encoded bridgeRNA molecule is programmable by modifying sequences in the target and / or donor binding loops of the bridgeRNA thereby engineering thebridgeRNA to specifically bind sequences of interest. In one category of programmable transposition, the bridgeRNA is designed such that a donor DNA molecule of interest can be recombined with a target DNA molecule of interest to effectuate insertion of a sequence located on a different DNA molecule or translocation of sequences on different DNA molecules. In another category of programmable transposition, the bridgeRNA is designed such that a donor DNA sequence of interest can be recombined with a target DNA of interest to effectuate excision or inversion of intervening sequences located on the same DNA molecule. Further, the invention also encompasses non-programmed uses of the IS110 family of transposons. For example, a non-programmed IS110 bridgeRNA (target and donor binding loops are not modified to change the binding specificity of the bridgeRNA) and transposase complex can be used to target naturally occurring target and donor site sequences in prokaryotic genomes, naturally occurring target and donor site sequences in eukaryotic genomes, introduced target and donor site sequences in prokaryotic genomes, and introduced target and donor site sequences in eukaryotic genomes.

[0173] A. Definitions

[0174] The terms “polynucleotide”, “nucleotide sequence”, “nucleic acid”, “nucleic acid molecule”, “nucleic acid segment”, and “oligonucleotide” are used interchangeably. They refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof. Polynucleotides may have any three-dimensional structure, and may perform any function, known or unknown. The following are non-limiting examples of polynucleotides: coding or non-coding regions of a gene or gene fragment, loci (locus) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short-hairpin RNA (shRNA), micro- RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. A polynucleotide may comprise one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. The sequence of nucleotides may be interrupted by non-nucleotide components. A polynucleotide may be further modified after polymerization, such as by conjugation with a labeling component.

[0175] The term “DNA” refers to, without limitation, deoxyribonucleic acid, which includes, but is not limited to genomic or non-genomic DNA that exists within a cell or the isolated form of such DNA. Genomic or non-genomic DNA includes without limitation,chromosomal or non-chromosomal DNA such as episomal, viral, plasmid, mitochondrial, cellular, or chloroplast DNA.

[0176] The terms “polypeptide”, “peptide” and “protein” are used interchangeably herein to refer to polymers of amino acids of any length. The polymer may be linear or branched, it may comprise modified amino acids, and it may be interrupted by non-amino acids. The terms also encompass an amino acid polymer that has been modified; for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation, such as conjugation with a labeling component. As used herein the term “amino acid” includes natural and / or unnatural or synthetic amino acids, including glycine and both the D or L optical isomers, and amino acid analogs and peptidomimetics.

[0177] As used herein, the term “nuclease” refers to an agent capable of breaking a phosphodiester bond linking a nucleotide residue in a nucleic acid molecule, such as a protein or small molecule. In some embodiments, the nuclease is an enzyme capable of binding a nucleic acid molecule, and breaking a phosphodiester bond that links a nucleotide residue within the nucleic acid molecule. The nuclease may be an endonuclease that cleaves the phosphodiester bond within the polynucleotide chain or an exonuclease that cleaves the phosphodiester bond at the end of the polynucleotide chain. In some embodiments, the nuclease is a site-specific nuclease that binds to and / or cleaves a particular phosphodiester bond within a particular nucleotide sequence. The term “nickase” refers to an endonuclease which cleaves only a single strand of a DNA duplex.

[0178] As used herein, the term “excisionase” refers to a host-derived, bacteriophage, or mobile genetic element sequence-specific DNA binding protein. It is involved in removing DNA from nucleotide sequences, repairing the DNA with or without a sequence scar. The removed DNA may be in the form of linear or circular ssDNA or dsDNA.

[0179] “Sequence-specific” refers to, but is not limited to, recombination or a recombination event which occurs at a predictable locus or identifiable nucleotide sequence or modification of a nucleotide at a predetermined sequence location.

[0180] The term “transposon”, as used herein, refers to a polynucleotide (or nucleic acid segment), which can be copied or moved to a new nucleic acid sequence context through the action of a transposase. An insertion sequence (IS) element refers to a transposon that encodes the minimal components necessary for recombination of a nucleotide sequence (e.g. a transposase and a bridgeRNA). An IS element may be referred to herein as an IS element, IS110 element, or transposon.

[0181] The term “transposase” as used herein refers to an enzyme, which is a component of a functional nucleic acid-protein complex (e.g., a transpososome) capable of transposition and which mediates transposition. The transposase may comprise a single protein or comprise multiple proteins. A transposase may be an enzyme capable of forming a functional complex with a transposon end, transposon end sequences, or transposon-derived sequences. The term “transposase” may also refer in certain embodiments to integrases, recombinases, invertases, or excisionases. Described herein are transposases derived from the IS110 family of transposons. The IS110 “transposases” described herein are also referred to as IS 110 “recombinases.”

[0182] The expression “transposition reaction” used herein refers to a reaction wherein a transposase recombines a DNA polynucleotide comprising a donor site sequence with a DNA polynucleotide comprising a target site sequence or with a DNA polynucleotide comprising a second donor site sequence, e.g., an integration reaction, a recombination reaction, an inversion reaction, or an excision reaction. A transposition reaction may occur when the donor site sequence and target site sequence are on one or more DNA molecules. A transposition reaction may occur when the donor site sequence and second donor site sequence are on one or more DNA molecules. Both the target, donor site, and second donor site sequences may contain a sequence or secondary structure. The target site, the donor site, and the second donor site sequences may contain a sequence or secondary structure recognized by the transposase and / or an insertion motif sequence where the transposase cuts or creates staggered breaks in the target polynucleotide sequence.

[0183] The term “transposon end sequence” as used herein refers to the nucleotide sequences at the distal ends of a transposon. The transposon end sequences, or subsequences thereof, may be the DNA sequences recognized by the transposase to form a transpososome complex and to perform a transposition reaction. In certain embodiments described herein, the transposon end sequences are derived from the non-coding end sequences of the IS110 family of transposons.

[0184] The term “near the 3' end” or “near the 5' end” of a sequence as used herein refers within 1, 2, 3, 4, or 5 nucleotides of the recited end of a nucleotide feature or component.

[0185] The practice of aspects of the present invention can employ, unless otherwise indicated, conventional techniques of cell biology, cell culture, molecular biology, transgenic biology, microbiology, recombinant DNA, and biochemistry, which are within the skill of theart. Such techniques are explained fully in the literature. See, e.g., Molecular Cloning A Laboratory Manual, 3rd Ed., ed. by Sambrook (2001), Fritsch and Maniatis (Cold Spring Harbor Laboratory Press: 1989); DNA Cloning, Volumes I and II (D. N. Glover ed., 1985); Oligonucleotide Synthesis (M. J. Gait ed., 1984); Mullis et al. U.S. Pat. No: 4,683,195; Nucleic Acid Hybridization (B. D. Hames & S. J. Higgins eds. 1984); Transcription and Translation (B. D. Hames & S. J. Higgins eds. 1984); Culture Of Animal Cells (R. I. Freshney, Alan R. Liss, Inc., 1987); Immobilized Cells and Enzymes (IRL Press, 1986); B. Perbal, A Practical Guide To Molecular Cloning (1984); the series, Methods In Enzymology (Academic Press, Inc., N.Y.), specifically, Methods In Enzymology, Vols. 154 and 155 (Wu et al. eds.); Gene Transfer Vectors For Mammalian Cells (J. H. Miller and M. P. Calos eds., 1987, Cold Spring Harbor Laboratory); Immunochemical Methods In Cell And Molecular Biology (Caner and Walker, eds., Academic Press, London, 1987); Handbook Of Experimental Immunology, Volumes I-FV (D. M. Weir and C. C. Blackwell, eds., 1986); Manipulating the Mouse Embryo, (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., 1986) and subsequent versions thereof.

[0186] One skilled in the art can obtain a protein in several ways, which include, but are not limited to, isolating the protein via biochemical means or expressing a nucleotide sequence encoding the protein of interest by genetic engineering methods.

[0187] A protein is encoded by a nucleic acid (including, for example, genomic DNA, messenger RNA (mRNA), complementary DNA (cDNA), synthetic DNA, as well as any form of corresponding RNA). Nucleic acids encoding a protein can be produced via recombinant DNA technology and such recombinant nucleic acids can be prepared by conventional techniques, including chemical synthesis, genetic engineering, enzymatic techniques, or a combination thereof.

[0188] B. IS110 Elements

[0189] The IS110 family of transposons refer to a family of transposons that are widespread in prokaryotic genomes. They are categorized into two groups, the IS110 group and the IS1111 group, and they encode transposases that cumulatively demonstrate a range of insertion site specificities. The IS110 transposases can exhibit invertase and excisionase activity, in addition to their transposase activity.

[0190] The life-cycle of an IS 110 element is depicted in Figure 1C. A linear IS110 element integrated into a target site comprises a left non-coding end (LE), a coding sequence for a transposase (Tpase) and a right non-coding end (RE), and, in some embodiments, theIS110 element is flanked by a repeated core sequence as shown in Figure 1C. IS110 elements excise themselves, resulting in pre-insertion (“target”) site bearing LF-core (if present)-RF, and a circular element with RE-core (if present)-LE-Tpase. Concatenation of the RE-LE junction forms a “donor” site sequence as a subsequence of the RE-LE junction, which, if present, includes the other core sequence found on the integrated element. The donor site sequence may also include sub-terminal inverted repeats (STIR). Concatenation of the RE-LE may also form a promoter which, in the appropriate cellular context, may promote expression from the LE or RE of a RNA molecule referred to herein as bridgeRNA. The promoter may also promote expression of the transposase in the appropriate cellular context. The bridgeRNA encoded within the LE or RE forms an RNA-protein complex with the transposase and recognizes the donor site and / or the target site sequences to mediate transposition. The circular form of the element can reinsert into the target site or insert into any other target site sequence recognized by the bridgeRNA-transposase complex.

[0191] The left non-coding end (LE) of an IS110 element refers to the nucleotide sequence 5' of the start codon of the IS110 element encoded IS110 transposase that extends (upstream) to the core or the 5' terminal end of the element. Thus, LE is simply defined as the sequence that comes between the CDS and the core or the 5' end of the element. The 5' terminal end of the element may be defined using comparative (meta)genomics, or by analysis of bridgeRNA specificity for the donor sequences found at the terminus of the LE or by BLAST similarity searches to IS 110s with terminal ends defined with the previous two methods. See Examples 1 and 2. In some embodiments, LE comprises an LE sequence provided in Figure 15 (SEQ ID NOS: 1-348) or Figure 17 (SEQ ID NOS: 30354-30529). In some embodiments, LE comprises an LE sequence provided in SEQ ID NOS: 349-10175 or 30530-40356.

[0192] The right non-coding end (RE) of an IS 110 element refers to the nucleotide sequence 3' of the stop codon of the IS110 element encoded IS110 transposase that extends (downstream) to the core or the 3' terminal end of the element. Thus, RE is simply defined as the sequence that comes between the CDS and the core or the 3' end of the element. The 3' terminal end of the element may be defined using comparative (meta)genomics or by analysis of bridgeRNA specificity for the donor sequences found at the terminus of the RE or by BLAST similarity searches to IS 110s with terminal ends defined with the previous two methods. See Examples 1 and 2. In some embodiments, RE comprises an RE sequence provided in Figure 15 (SEQ ID NOS: 1-348) or Figure 17 (SEQ ID NOS: 30354-30529). Insome embodiments, RE comprises an RE sequence provided in SEQ ID NOS: 349-10175 or 30530-40356.

[0193] For IS110 transposons that comprise a core sequence, the core refers to an identical nucleotide sequence found immediately 5' and 3' of the left non-coding end (LE) and right non-coding end (RE), respectively. The core was previously referred to as “target intervening core” or “TIC” and any references to target intervening core or TIC refer to the core sequence. In some embodiments, the core sequence is 1-10 nucleotides long. In some embodiments, the core sequence is 1-5 nucleotides long. In some embodiments, the core sequence is 1 nucleotide long, 2 nucleotides long, 3 nucleotides long, 4 nucleotides long, 5 nucleotides long, 6 nucleotides long, 7 nucleotides long, 8 nucleotides long, 9 nucleotides long, or 10 nucleotides long. In certain embodiments, the core sequence is 2 nucleotides long. In some embodiments, a core comprises a core sequence provided in Figure 15 (SEQ ID NOS: 1-348) or Figure 17 (SEQ ID NOS: 30354-30529). In some embodiments, a core comprises a core sequence provided in SEQ ID NOS: 349-10175 or 30530-40356.

[0194] Exemplary IS110 family IS element sequences are provided in Figure 15 (SEQ ID NOS: 1-348). The nucleotide sequences of LE, core (where present), the transposase, and RE are indicated as described above for Figure 15. Additional exemplary IS110 family IS element sequences are provided in SEQ ID NOS: 349-10175.

[0195] For IS110 elements that comprise a core sequence, RE-core-LE refers to a concatenation of the nucleotide sequences of the RE, core, and LE which a portion thereof (e.g., the donor site sequence comprised of LD-core-RD) may be bound by an IS110 family transposase described herein (e.g., see Section C). In some embodiments, RE-core-LE comprises an LE, core, and RE provided in Figure 15 (SEQ ID NOS: 1-348) or Figure 17 (SEQ ID NOS: 30354-30529). In some embodiments, RE-core-LE comprises an LE, core, and RE provided in SEQ ID NOS: 349-10175 or 30530-40356. In some embodiments, LD- core-RD comprises an LD, core, and RD provided in Figure 17 (SEQ ID NOS: 30354- 30529). The nucleotide sequences of LD and RD are indicated in bold and the nucleotide sequence of core sequence is represented as non-highlighted text with a single underline. In some embodiments, LD-core-RD comprises an LD, core, and RD derived from the LDG and RDG provided in Figure 15 (SEQ ID NOS: 1-348).

[0196] For IS110 elements that comprise a core sequence, LF-core-RF refers to a concatenation of the nucleotide sequences of the LF, core, and RF which a portion thereof (e.g., the target site sequence comprised of LT-core-RT) may be bound by an IS110 familytransposase described herein (e.g., see Section C). In some embodiments, LF-core-RF comprises a LF, core, and RF provided in Figure 18 (SEQ ID NOs: 20351-20526). In some embodiments, RF-core-LF comprises an LF, core, and RF provided in SEQ ID NOs: 20527- 30353. In some embodiments, LT-core-RT comprises an LT, core, and RT provided in Figure 18 (SEQ ID NOS: 20351-20368). The nucleotide sequences of LT and RT are indicated in bold and the nucleotide sequence of core sequence is represented as nonhighlighted text with a single underline. In some embodiments, LT-core-RT comprises an LT, core, and RT derived from the LTG and RTG provided in Figure 15 (SEQ ID NOS: 1- 348).

[0197] For IS110 elements that do not comprise a core sequence, RE-LE refers to a concatenation of the nucleotide sequences of the RE and LE which a portion thereof (e.g., the donor site sequence comprises of LD-RD) may be bound by an IS 110 family transposase described herein (e.g., see Section C). In some embodiments, RE-LE comprises an LE and RE provided in Figure 15 (SEQ ID NOS: 1-348) or Figure 17 (SEQ ID NOS: 30354-30529). In some embodiments, RE-LE comprises an LE and RE provided in SEQ ID NOS: 349- 10175 or 30530-40356. In some embodiments, LD-RD comprises an LD and RD derived from the LDG and RDG provided in Figure 15 (SEQ ID NOS: 1-348).

[0198] For IS110 elements that do not comprise a core sequence, LF-RF refers to a concatenation of the nucleotide sequences of the LF and RF which a portion thereof (e.g., the target site sequence comprised of LT-RT) may be bound by an IS110 family transposase described herein (e.g., see Section C). In some embodiments, LF-RF comprises an LF and RF provided in Figure 18 (SEQ ID NOs: 20351-20526). In some embodiments, LF-RF comprises a LF and RF provided in SEQ ID NOs: 20527-30353. In some embodiments, LT- RT comprises an LT and RT derived from the LTG and RTG provided in Figure 15 (SEQ ID NOS: 1-348).

[0199] C. IS110 Family Transposases

[0200] The IS110 family of transposases encoded within the IS110 transposons were identified by homology searching for a DEDD catalytic domain, which is a RuvC-like domain. See Example 1. IS110 family transposases described herein comprise an N-terminal RuvC-like DEDD catalytic domain and a C-terminal transposase domain with two canonical Pfam domains as depicted in Figure IB. In some embodiments, the polypeptide sequence between the N-terminal RuvC-like DEDD catalytic domain and the C-terminal transposase domain comprises a linker domain comprising a coiled-coil.

[0201] Within the IS110 family, transposons can be classed into the IS110 group which are any insertion sequence (IS) element encoding an IS110 transposase and comprising a longer 5' non-coding end (LE) than 3' non-coding end (RE). See FIGS. 1A, D, E.

[0202] Within the IS110 family, transposons can be classed into the IS1111 group which are any insertion sequence (IS) element encoding an IS110 transposase and typically comprising a longer 3 'non-coding end (RE) than 5' non-coding end (LE). See FIGS. 1A, D, E.

[0203] Exemplary primary amino acid sequences and secondary structure prediction of IS 110 family transposases are provided in Figure 16 (SEQ ID NOS: 10176-10523).Additional exemplary primary amino acid sequences of IS 110 family transposases are provided in SEQ ID NOS: 10524-20350. The polypeptide sequence between the N-terminal RuvC-like DEDD catalytic domain and the C-terminal transposase domain comprises a linker domain comprising a coiled-coil.

[0204] In some embodiments, the IS110 family transposase comprises an amino acid sequence that is 50% identical or more, 51% identical or more, 52% identical or more, 53% identical or more, 54% identical or more, 55% identical or more, 56% identical or more, 57% identical or more, 58% identical or more, 59% identical or more, 60% identical or more, 61% identical or more, 62% identical or more, 63% identical or more, 64% identical or more, 65% identical or more, 66% identical or more, 67% identical or more, 68% identical or more, 69% identical or more, 70% identical or more, 71% identical or more, 72% identical or more, 73% identical or more, 74% identical or more, 75% identical or more, 76% identical or more, 77% identical or more, 78% identical or more, 79% identical or more, 80% identical or more, 81% identical or more, 82% identical or more, 83% identical or more, 84% identical or more, 85% identical or more, 86% identical or more, 87% identical or more, 88% identical or more, 89% identical or more, 90% identical or more, 91% identical or more, 92% identical or more, 93% identical or more, 94% identical or more, 95% identical or more, 96% identical or more, 97% identical or more, 98% identical or more, 99% identical or more, or 100% identical to a sequence provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524- 20350. In some embodiments the sequence is “protein_IS621” of Figure 16 (SEQ ID NO: 10176). In some embodiments, the sequence is “protein_IS1621_23122” (SEQ ID NO: 10181), also referred to herein as “23122”. In some embodiments, the IS110 family transposases for use in human cell applications is “protein_IS1621_23122” (SEQ ID NO: 10181), also referred to herein as “23122”. See e.g., Figures 24, 25, 26, 27, 28, 62, 63, 64, 65, 66, 67, 68, 69, 70, 73, 74, 75, 76, and 77 which use protein_IS1621_23122.

[0205] Domain motifs and / or regions of the IS110 family transposase can be identified not necessarily by similarity of amino acid sequences but by structural similarity. In some embodiments, structural similarity is determined by the template modeling score (TM-score). See Example 15 and Figures 14A-H. In some embodiments, predicted secondary structure is used to identify domain motifs and / or regions of the IS110 family transposase. In some embodiments, secondary structure of a primary amino acid sequence is predicted using a standard mkdssp tool on tertiary structure files or equivalent protein structure prediction software. In some embodiments, the linker domain of the IS110 family transposase comprises a polypeptide sequence between the RuvC-like DEDD catalytic domain and transposase domain that comprises an amino acid sequence that is predicted to form a coiled-coil.

[0206] In some embodiments, the IS110 family transposase comprises a polypeptide that forms a similar tertiary structure to the tertiary structure of IS621 as shown in Figure 14C. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to that of IS621 if the template modeling score (TM-score) is 0.5 or higher. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to that of IS621 if the template modeling score (TM-score) is 0.5 or higher, 0.6 or higher, 0.7 or higher, 0.8 or higher, or 0.9 or higher.

[0207] The TM-score is defined as provided in Figure 14A, where Ltarget is the length of the amino acid sequence of the target protein, and Lcommon is the number of residues that appear in both the template and target structures, di is the distance between the ith pair of residues in the template and target structures, and do(Ltarget) = 1 ,24-^(L_target-l 5)-l .8 is a distance scale that normalizes distances (Zhang, Yang, and Jeffrey Skolnick. 2004. “Scoring Function for Automated Assessment of Protein Structure Template Quality.” Proteins 57 (4): 702-10). Alternatively, the score can be normalized according to the length of the query protein, or the score can be normalized by the averaged length of the two proteins. A TM- score has a value in (0,1], and a cutoff of >0.5 is commonly used for identifying proteins with homologous tertiary structures (Zhang, Yang, and Jeffrey Skolnick. 2005. “TM-Align: A Protein Structure Alignment Algorithm Based on the TM-Score.” Nucleic Acids Research 33 (7): 2302-9).

[0208] In some embodiments, the IS110 family transposase comprises an amino acidsequence that is 15% identical or more, 16% identical or more, 17% identical or more, 18% identical or more, 19% identical or more, 20% identical or more, 21% identical or more, 22% identical or more, 23% identical or more, 24% identical or more, 25% identical or more, 26% identical or more, 27% identical or more, 28% identical or more, 29% identical or more, 30% identical or more, 31% identical or more, 32% identical or more, 33% identical or more, 34% identical or more, 35% identical or more, 36% identical or more, 37% identical or more, 38% identical or more, 39% identical or more, 40% identical or more, 41% identical or more, 42% identical or more, 43% identical or more, 44% identical or more, 45% identical or more, 46% identical or more, 47% identical or more, 48% identical or more, 49% identical or more, 50% identical or more, 51% identical or more, 52% identical or more, 53% identical or more, 54% identical or more, 55% identical or more, 56% identical or more, 57% identical or more, 58% identical or more, 59% identical or more, 60% identical or more, 61% identical or more, 62% identical or more, 63% identical or more, 64% identical or more, 65% identical or more, 66% identical or more, 67% identical or more, 68% identical or more, 69% identical or more, 70% identical or more, 71% identical or more, 72% identical or more, 73% identical or more, 74% identical or more, 75% identical or more, 76% identical or more, 77% identical or more, 78% identical or more, 79% identical or more, 80% identical or more, 81% identical or more, 82% identical or more, 83% identical or more, 84% identical or more, 85% identical or more, 86% identical or more, 87% identical or more, 88% identical or more, 89% identical or more, 90% identical or more, 91% identical or more, 92% identical or more, 93% identical or more, 94% identical or more, 95% identical or more, 96% identical or more, 97% identical or more, 98% identical or more, 99% identical or more, or 100% identical to a sequence provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350 and forms a similar tertiary structure as shown in Figure 14C. In some embodiments the sequence is “protein_IS621” of Figure 16 (SEQ ID NO: 10176). In some embodiments, the sequence is “protein_IS1621_23122” of Figure 16 (SEQ ID NO: 10181). In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to that of IS621 if the template modeling score (TM-score) is 0.5 or higher. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to that of IS621 if the template modeling score (TM-score) is 0.5 or higher, 0.6 or higher, 0.7 or higher, 0.8 or higher, or 0.9 or higher.

[0209] In certain aspects, described herein is an IS110 family transposase comprisingmeans for performing a transposase reaction. In some embodiments, the means for performing a transposase reaction comprises a sequence provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350.

[0210] In certain aspects, described herein are nucleic acids encoding any of the IS110 family transposase amino acid sequences provided herein.

[0211] C.l. RuvC-like DEDD Catalytic Domain

[0212] The RuvC-like DEDD catalytic domain refers to the domain of the IS110 transposase that resembles the RuvC Holliday junction resolvase, an abundant protein domain found within proteins of diverse function. The RuvC domain is often found within RNA- guided CRISPR nucleases. RNA-guided RuvC domain bearing CRISPR nucleases are sometimes associated with transposons, such as CRISPR associated transposons (CAST).CRISPR nucleases associated with transposons do not mediate transposition but impart target specificity for the transposome. The IS110 family transposases described herein comprise a RuvC-like DEDD catalytic domain.

[0213] In some embodiments, the IS110 family transposase comprises a RuvC-like DEDD catalytic domain comprising an amino acid sequence that is 50% identical or more, 51% identical or more, 52% identical or more, 53% identical or more, 54% identical or more,55% identical or more, 56% identical or more, 57% identical or more, 58% identical or more,59% identical or more, 60% identical or more, 61% identical or more, 62% identical or more,63% identical or more, 64% identical or more, 65% identical or more, 66% identical or more,67% identical or more, 68% identical or more, 69% identical or more, 70% identical or more,71% identical or more, 72% identical or more, 73% identical or more, 74% identical or more,75% identical or more, 76% identical or more, 77% identical or more, 78% identical or more,79% identical or more, 80% identical or more, 81% identical or more, 82% identical or more,83% identical or more, 84% identical or more, 85% identical or more, 86% identical or more,87% identical or more, 88% identical or more, 89% identical or more, 90% identical or more,91% identical or more, 92% identical or more, 93% identical or more, 94% identical or more,95% identical or more, 96% identical or more, 97% identical or more, 98% identical or more,99% identical or more, or 100% identical to the RuvC-like DEDD catalytic domain sequences provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524- 20350. In some embodiments, the RuvC-like DEDD catalytic domain sequence is the RuvC- like DEDD catalytic domain of “protein_IS621” of Figure 16. In some embodiments, the RuvC-like DEDD catalytic domain sequence is the RuvC-like DEDD catalytic domain of“protein_IS1621_23122” of Figure 16.

[0214] In some embodiments, the IS110 family transposase comprises a RuvC-like DEDD catalytic domain that forms a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the RuvC- like DEDD catalytic domain of IS621 if the template modeling score (TM-score) for the RuvC-like DEDD catalytic domain is 0.5 or higher. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621 if the template modeling score (TM-score) is 0.5 or higher, 0.6 or higher, 0.7 or higher, 0.8 or higher, or 0.9 or higher.

[0215] In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621 based on distances between the alpha carbon of conserved residues in the RuvC-like DEDD catalytic domain. Figure 16 (SEQ ID NOS: 10176-10523) provides in bold typeface up to 5 amino acids that are highly conserved in the RuvC-like DEDD catalytic domain. Conserved amino acids in a particular amino acid sequence are identified by primary amino acid sequence alignment. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621 if a distance (“DI”) between the alpha carbon of a first conserved residue and the alpha carbon of a second conserved residue of the amino acid sequence is less than 10 angstroms (A), wherein the conserved residues of the amino acid sequence is per alignment of the primary amino acid sequence with one or more RuvC-like DEDD catalytic domains, such as IS621. In IS621 the first conserved residue is DI 1 and the second conserved residue is E60. In some embodiments, DI is between 4 and 10 angstroms (A). In some embodiments, DI is between 5 and 7.5 angstroms (A). In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621 if an average distance (“D2”) between the alpha carbon of a first conserved residue and the alpha carbon of a third conserved residue, between the alpha carbon of a first conserved residue and the alpha carbon of a fourth conserved residue, and between the alpha carbon of a first conserved residue and the alpha carbon of a fifth conserved residue, is less than 10 angstroms (A), wherein the conserved residues of the amino acid sequence is per alignment of the primary amino acid sequence with one or more RuvC-like DEDD catalytic domains,such as IS621. In IS621 the first conserved residue is DI 1, the third conserved residue is K100, the fourth conserved residue is D102, and the fifth conserved residue is D105. In some embodiments, D2 is between 5 and 10 angstroms (A). In some embodiments, D2 is between 7.5 and 10 angstroms (A). In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621 if an average distance (“D3”) between the alpha carbon of a second conserved residue and the alpha carbon of a third conserved residue, between the alpha carbon of a second conserved residue and the alpha carbon of a fourth conserved residue, and between the alpha carbon of a second conserved residue and the alpha carbon of a fifth conserved residue, is less than 15 angstroms (A), wherein the conserved residues of the amino acid sequence is per alignment of the primary amino acid sequence with one or more RuvC-like DEDD catalytic domains, such as IS621. In IS621 the second conserved residue is E60, the third conserved residue is KI 00, the fourth conserved residue is DI 02, and the fifth conserved residue is DI 05. In some embodiments, D3 is between 10 and 15 angstroms (A). In some embodiments, D3 is between 13 and 15 angstroms (A). In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621 if DI is less than 10 angstroms (A), D2 is less than 10 angstroms (A), and D3 is less than 15 angstroms (A). In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621 if DI is between 4 and 10 angstroms (A), D2 is between 5 and 10 angstroms (A), and D3 is between 10 and 15 angstroms (A). In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621 if DI is between 5 and 7.5 angstroms (A), D2 is between 7.5 and 10 angstroms (A), and D3 is between 13 and 15 angstroms (A).

[0216] In some embodiments, the IS110 family transposase comprises a RuvC-like DEDD catalytic domain comprising an amino acid sequence that is 15% identical or more, 16% identical or more, 17% identical or more, 18% identical or more, 19% identical or more,20% identical or more, 21% identical or more, 22% identical or more, 23% identical or more,24% identical or more, 25% identical or more, 26% identical or more, 27% identical or more,28% identical or more, 29% identical or more, 30% identical or more, 31% identical or more,32% identical or more, 33% identical or more, 34% identical or more, 35% identical or more,36% identical or more, 37% identical or more, 38% identical or more, 39% identical or more,40% identical or more, 41% identical or more, 42% identical or more, 43% identical or more,44% identical or more, 45% identical or more, 46% identical or more, 47% identical or more,48% identical or more, 49% identical or more, 50% identical or more, 51% identical or more,52% identical or more, 53% identical or more, 54% identical or more, 55% identical or more,56% identical or more, 57% identical or more, 58% identical or more, 59% identical or more,60% identical or more, 61% identical or more, 62% identical or more, 63% identical or more,64% identical or more, 65% identical or more, 66% identical or more, 67% identical or more,68% identical or more, 69% identical or more, 70% identical or more, 71% identical or more,72% identical or more, 73% identical or more, 74% identical or more, 75% identical or more,76% identical or more, 77% identical or more, 78% identical or more, 79% identical or more,80% identical or more, 81% identical or more, 82% identical or more, 83% identical or more,84% identical or more, 85% identical or more, 86% identical or more, 87% identical or more,88% identical or more, 89% identical or more, 90% identical or more, 91% identical or more,92% identical or more, 93% identical or more, 94% identical or more, 95% identical or more,96% identical or more, 97% identical or more, 98% identical or more, 99% identical or more, or 100% identical to the RuvC-like DEDD catalytic domain sequences provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350 and forms a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621 if the template modeling score (TM-score) for the RuvC-like DEDD catalytic domain is 0.5 or higher. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621 if the template modeling score (TM-score) is 0.5 or higher, 0.6 or higher, 0.7 or higher, 0.8 or higher, or 0.9 or higher. In some embodiments the RuvC-like DEDD catalytic domain sequence provided is the RuvC- like DEDD catalytic domain of “protein_IS621” of Figure 16 (SEQ ID NO: 10176). In some embodiments the RuvC-like DEDD catalytic domain sequence provided is the RuvC-like DEDD catalytic domain of “protein_IS1621_23122” of Figure 16 (SEQ ID NO:10181).

[0217] In some embodiments, the RuvC-like DEDD catalytic domain can be identified using statistical models that annotate protein domains, such as Pfam profile hidden markov models (pHMMs). In IS 110 family transposases, the RuvC-like DEDD catalytic domain is often recognized by the Pfam profile hidden markov model (pHMM) PF01548, short name DEDD Tnp ISl 10.

[0218] In some embodiments, the IS110 family transposase comprises a RuvC-like DEDD catalytic domain comprising a motif D-x(43)-E-x(39)-K-x(l)-D-x(2)-D, D-x(42)-E- x(34)-K-x(l)-D-x(2)-D, [DE]-x(38,63)-[EACDGQVIPS]-x(30,53)-[KSQIVRHLTMA]-x(l)- [DNE]-x(2)-[DEASCM], [DE]-x(41,59)-[EYALGHCVFITMS]-x(30,45)-[KRMQNH]-x(l)- [DN]-x(2)-[DAS], GIDVS, GLDVH,[GAS] [ILVDFMACWTHGN] [DE] [VTIWPLAFCSMR] [S AHGCD], [GAS][LIMVCFA]D[VLIFAYQTHMDCWRGS][HASGD], MEATG, MEACG, [MLVIFCYAGTWSHPR]-x(0,2)- [EGACQDVIMP][ASGPHIDRYLNQCTVFMEK][TSECPAYGIVDNFKLMHR]-x(O,l)- [GASTRDQVNLWEKYHICMP], [MYVCLIASFTGQEHW]-x(0,2)- [EYQGVACFLHITMSDKW]-x(0,2)-[AEVYSFMGTICNPLDQ]-x(0,2)- [CGATESMVLINPFDQW]-x(0,2)-[GPASTLCYINRMVFEWHQDK], DRIDA, DRRDA, [DNE] [RPAKVTSQEFGIDLHMNYWC] [ILKVAGTNRSFQPHDEMCWY] [D AESCM] [A CSPLTGV], or[DN][RAKEYDQGFVPTMSLHWNIC] [RALNIVKHTQDSEMGWF YCP] [DS A] [ASTGCV LI], with motifs in common Prosite format, where x is any amino acid and x(n) represents n number of any amino acid and x(n,m) represents n to m number of any amino acids. In some embodiments, the RuvC-like DEDD catalytic domain comprises a motif D-x(43)-E-x(39)-K- x(l)-D-x(2)-D, D-x(42)-E-x(34)-K-x(l)-D-x(2)-D, GIDVS, GLDVH, MEATG, MEACG, DRIDA, or DRRDA. In some embodiments, the RuvC-like DEDD catalytic domain comprising a motif above forms a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621 if the template modeling score (TM-score) for the RuvC-like DEDD catalytic domain is 0.5 or higher. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621 if the template modeling score (TM-score) is 0.5 or higher, 0.6 or higher, 0.7 or higher, 0.8 or higher, or 0.9 or higher.

[0219] In some embodiments, the IS110 family transposase comprises a RuvC-like DEDD catalytic domain comprising an amino acid sequence that is 50% identical or more, 51% identical or more, 52% identical or more, 53% identical or more, 54% identical or more, 55% identical or more, 56% identical or more, 57% identical or more, 58% identical or more,59% identical or more, 60% identical or more, 61% identical or more, 62% identical or more,63% identical or more, 64% identical or more, 65% identical or more, 66% identical or more,67% identical or more, 68% identical or more, 69% identical or more, 70% identical or more,71% identical or more, 72% identical or more, 73% identical or more, 74% identical or more,75% identical or more, 76% identical or more, 77% identical or more, 78% identical or more,79% identical or more, 80% identical or more, 81% identical or more, 82% identical or more,83% identical or more, 84% identical or more, 85% identical or more, 86% identical or more,87% identical or more, 88% identical or more, 89% identical or more, 90% identical or more,91% identical or more, 92% identical or more, 93% identical or more, 94% identical or more,95% identical or more, 96% identical or more, 97% identical or more, 98% identical or more,99% identical or more, or 100% identical to a RuvC-like DEDD catalytic domain sequence of Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350 and comprising any of the motifs in the preceding paragraph. In some embodiments, the amino acid sequence of the IS110 family transposase comprises a RuvC-like DEDD catalytic domain forms a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621 if the template modeling score (TM-score) for the RuvC-like DEDD catalytic domain is 0.5 or higher. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621 if the template modeling score (TM-score) is 0.5 or higher, 0.6 or higher, 0.7 or higher, 0.8 or higher, or 0.9 or higher.

[0220] The present invention contemplates domain swapping in order to generate IS110 transposase chimeras that have advantageous functions. In some embodiments, exchanging of a RuvC-like DEDD catalytic domain (e.g., any RuvC-like DEDD catalytic domain provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350) with a different RuvC-like DEDD catalytic domain (e.g., any other RuvC-like DEDD catalytic domain provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350) resulting in advantageous properties is also envisioned as described in Farruggio et al., 2014. For example, exchanging one RuvC-like DEDD catalytic domain with another RuvC-like DEDD catalytic domain may allow for higher affinity for a bridgeRNA or increased transposition efficiency.

[0221] C.2. Transposase Domain

[0222] The IS110 family transposases described herein comprise a transposase domain.

[0223] In some embodiments, the transposase domain can be identified using statistical models that annotate protein domains, such as Pfam profile hidden markov models (pHMMs). In IS 110 family transposases, the transposase domain is often recognized by the Pfam profile hidden markov model (pHMM) PF02371, short name Transposase_20.

[0224] In some embodiments, the IS110 family transposase comprises a transposase domain comprising an amino acid sequence that is 50% identical or more, 51% identical or more, 52% identical or more, 53% identical or more, 54% identical or more, 55% identical or more, 56% identical or more, 57% identical or more, 58% identical or more, 59% identical or more, 60% identical or more, 61% identical or more, 62% identical or more, 63% identical or more, 64% identical or more, 65% identical or more, 66% identical or more, 67% identical or more, 68% identical or more, 69% identical or more, 70% identical or more, 71% identical or more, 72% identical or more, 73% identical or more, 74% identical or more, 75% identical or more, 76% identical or more, 77% identical or more, 78% identical or more, 79% identical or more, 80% identical or more, 81% identical or more, 82% identical or more, 83% identical or more, 84% identical or more, 85% identical or more, 86% identical or more, 87% identical or more, 88% identical or more, 89% identical or more, 90% identical or more, 91% identical or more, 92% identical or more, 93% identical or more, 94% identical or more, 95% identical or more, 96% identical or more, 97% identical or more, 98% identical or more, 99% identical or more, or 100% identical to the transposase domain sequences provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350. In some embodiments the transposase domain sequence provided is the transposase domain of “protein_IS621” of Figure 16 (SEQ ID NO: 10176). In some embodiments the transposase domain sequence provided is the transposase domain of “protein_IS1621_23122” (SEQ ID NO: 10181).

[0225] In some embodiments, the IS110 family transposase comprises a transposase domain that forms a similar tertiary structure to the transposase domain of IS621. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 if the template modeling score (TM-score) for the transposase domain is 0.5 or higher. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 if the template modeling score (TM-score) is 0.5 or higher, 0.6or higher, 0.7 or higher, 0.8 or higher, or 0.9 or higher.

[0226] In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 based on distances between the alpha carbon of conserved residues in the transposase domain. Figure 16 (SEQ ID NOS: 10176-10523) provides in bold typeface amino acids up to 5 amino acids that are highly conserved in the transposase domain. Conserved amino acids in a particular amino acid sequence are identified by primary amino acid sequence alignment. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 if an average distance (“DI”) between the alpha carbon of a first conserved residues and the alpha carbon of a second conserved residue and between the alpha carbon of a first conserved residue and the alpha carbon of a fifth conserved residue of the amino acid sequence is less than 25 angstroms (A), wherein the conserved residues of the amino acid sequence is per alignment of the primary amino acid sequence with one or more transposase domains, such as IS621. In IS621 the first conserved residue is G203, the second conserved residue is G233, and the fifth conserved residue is G255. In some embodiments, DI is between 15 and 25 angstroms (A). In some embodiments, DI is between 17 and 23 angstroms (A). In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 if an average distance (“D2”) between the alpha carbon of a second conserved residue and the alpha carbon of a third conserved residue, between the alpha carbon of a second conserved residue and the alpha carbon of a fourth conserved residue, between the alpha carbon of a fifth conserved residue and the alpha carbon of a third conserved residue, and between the alpha carbon of a fifth conserved residue and the alpha carbon of a fourth conserved residue, is less than 25 angstroms (A), wherein the conserved residues of the amino acid sequence is per alignment of the primary amino acid sequence with one or more transposase domains, such as IS621. In IS621 the second conserved residue is G233, the third conserved residue is S241, the fourth conserved residue is G242, and the fifth conserved residue is G255. In some embodiments, D2 is between 20 and 25 angstroms (A). In some embodiments, D2 is between 22 and 24 angstroms (A). In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase of IS621 if a distance (“D3”) between the alpha carbon of a second conserved residue and the alpha carbon of a fifth conserved residue is less than 15 angstroms (A), wherein the conserved residues of the amino acid sequence is per alignment of the primary amino acid sequence with one or more transposase domains,such as IS621. In IS621 the second conserved residue is G233 and the fifth conserved residue is G255. In some embodiments, D3 is between 5 and 15 angstroms (A). In some embodiments, D3 is between 7 and 12 angstroms (A). In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 if DI is less than 25 angstroms (A), D2 is less than 25 angstroms (A), and D3 is less than 15 angstroms (A). In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 if DI is between 15 and 25 angstroms (A), D2 is between 20 and 25 angstroms (A), and D3 is between 5 and 15 angstroms (A). In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 if DI is between 17 and 23 angstroms (A), D2 is between 22 and 24 angstroms (A), and D3 is between 7 and 12 angstroms (A).

[0227] In some embodiments, the IS110 family transposase comprises a transposase domain comprising an amino acid sequence that is 15% identical or more, 16% identical or more, 17% identical or more, 18% identical or more, 19% identical or more, 20% identical or more, 21% identical or more, 22% identical or more, 23% identical or more, 24% identical or more, 25% identical or more, 26% identical or more, 27% identical or more, 28% identical or more, 29% identical or more, 30% identical or more, 31% identical or more, 32% identical or more, 33% identical or more, 34% identical or more, 35% identical or more, 36% identical or more, 37% identical or more, 38% identical or more, 39% identical or more, 40% identical or more, 41% identical or more, 42% identical or more, 43% identical or more, 44% identical or more, 45% identical or more, 46% identical or more, 47% identical or more, 48% identical or more, 49% identical or more, 50% identical or more, 51% identical or more, 52% identical or more, 53% identical or more, 54% identical or more, 55% identical or more, 56% identical or more, 57% identical or more, 58% identical or more, 59% identical or more, 60% identical or more, 61% identical or more, 62% identical or more, 63% identical or more, 64% identical or more, 65% identical or more, 66% identical or more, 67% identical or more, 68% identical or more, 69% identical or more, 70% identical or more, 71% identical or more, 72% identical or more, 73% identical or more, 74% identical or more, 75% identical or more, 76% identical or more, 77% identical or more, 78% identical or more, 79% identical or more, 80% identical or more, 81% identical or more, 82% identical or more, 83% identical or more, 84% identical or more, 85% identical or more, 86% identical or more, 87% identical or more, 88% identical or more, 89% identical or more, 90% identical or more, 91% identical or more, 92% identical ormore, 93% identical or more, 94% identical or more, 95% identical or more, 96% identical or more, 97% identical or more, 98% identical or more, 99% identical or more, or 100% identical to the transposase domain sequences provided in Figure 16 (SEQ ID NOS: 10176- 10523) or SEQ ID NOS: 10524-20350 and the amino acid sequence forms a similar tertiary structure to the transposase domain of IS621. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 if the template modeling score (TM-score) for the transposase domain is 0.5 or higher. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 if the template modeling score (TM-score) is 0.5 or higher, 0.6 or higher, 0.7 or higher, 0.8 or higher, or 0.9 or higher. In some embodiments the transposase domain sequence provided is the transposase domain of “protein_IS621” of Figure 16 (SEQ ID NO: 10176). In some embodiments the transposase domain sequence provided is the transposase domain of “protein_IS1621_23122” of Figure 16 (SEQ ID NO: 10181).

[0228] In some embodiments, the IS110 family transposase comprises a RuvC-like DEDD catalytic domain described in the preceding paragraphs (see Section C. l.) and further comprises a transposase domain comprising an amino acid sequence that is 50% identical or more, 51% identical or more, 52% identical or more, 53% identical or more, 54% identical or more, 55% identical or more, 56% identical or more, 57% identical or more, 58% identical or more, 59% identical or more, 60% identical or more, 61% identical or more, 62% identical or more, 63% identical or more, 64% identical or more, 65% identical or more, 66% identical or more, 67% identical or more, 68% identical or more, 69% identical or more, 70% identical or more, 71% identical or more, 72% identical or more, 73% identical or more, 74% identical or more, 75% identical or more, 76% identical or more, 77% identical or more, 78% identical or more, 79% identical or more, 80% identical or more, 81% identical or more, 82% identical or more, 83% identical or more, 84% identical or more, 85% identical or more, 86% identical or more, 87% identical or more, 88% identical or more, 89% identical or more, 90% identical or more, 91% identical or more, 92% identical or more, 93% identical or more, 94% identical or more, 95% identical or more, 96% identical or more, 97% identical or more, 98% identical or more, 99% identical or more, or 100% identical to the transposase domain sequences provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350. In some embodiments, the amino acid sequence forms a similar tertiary structure to the transposasedomain of IS621. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 if the template modeling score (TM-score) for the transposase domain is 0.5 or higher. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 if the template modeling score (TM- score) is 0.5 or higher, 0.6 or higher, 0.7 or higher, 0.8 or higher, or 0.9 or higher. In some embodiments the transposase domain sequence provided is the transposase domain of “protein_IS621” of Figure 16 (SEQ ID NO: 10176). In some embodiments the transposase domain sequence provided is the transposase domain of “protein_IS1621_23122” of Figure 16 (SEQ ID NO: 10181).

[0229] In some embodiments, the IS110 family transposase comprises a RuvC-like DEDD catalytic domain described in the preceding paragraphs (see Section C. l.) and further comprises a transposase domain that the amino acid sequence forms a similar tertiary structure to the transposase domain of IS621. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 if the template modeling score (TM-score) for the transposase domain is 0.5 or higher. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 if the template modeling score (TM-score) is 0.5 or higher, 0.6 or higher, 0.7 or higher, 0.8 or higher, or 0.9 or higher. In some embodiments the transposase domain sequence provided is the transposase domain of “protein_IS621” of Figure 16 (SEQ ID NO: 10176). In some embodiments the transposase domain sequence provided is the transposase domain of “protein_IS1621_23122” of Figure 16 (SEQ ID NO: 10181). In some embodiments, the amino acid sequence forms a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621 if the template modeling score (TM-score) for the RuvC-like DEDD catalytic domain is 0.5 or higher. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621 if the template modeling score (TM-score) is 0.5 or higher, 0.6 or higher,0.7 or higher, 0.8 or higher, or 0.9 or higher. In some embodiments the RuvC-hke DEDD catalytic domain sequence provided is the RuvC-like DEDD catalytic domain of “protein_IS621” of Figure 16 (SEQ ID NO: 10176). In some embodiments the RuvC-like DEDD catalytic domain sequence provided is the RuvC-like DEDD catalytic domain of “protein_IS1621_23122” of Figure 16 (SEQ ID NO: 10181). In some embodiments, the transposase domain comprises an amino acid sequence that is 15% identical or more, 16% identical or more, 17% identical or more, 18% identical or more, 19% identical or more, 20% identical or more, 21% identical or more, 22% identical or more, 23% identical or more, 24% identical or more, 25% identical or more, 26% identical or more, 27% identical or more, 28% identical or more, 29% identical or more, 30% identical or more, 31% identical or more, 32% identical or more, 33% identical or more, 34% identical or more, 35% identical or more, 36% identical or more, 37% identical or more, 38% identical or more, 39% identical or more, 40% identical or more, 41% identical or more, 42% identical or more, 43% identical or more, 44% identical or more, 45% identical or more, 46% identical or more, 47% identical or more, 48% identical or more, 49% identical or more, 50% identical or more, 51% identical or more, 52% identical or more, 53% identical or more, 54% identical or more, 55% identical or more, 56% identical or more, 57% identical or more, 58% identical or more, 59% identical or more, 60% identical or more, 61% identical or more, 62% identical or more, 63% identical or more, 64% identical or more, 65% identical or more, 66% identical or more, 67% identical or more, 68% identical or more, 69% identical or more, 70% identical or more, 71% identical or more, 72% identical or more, 73% identical or more, 74% identical or more, 75% identical or more, 76% identical or more, 77% identical or more, 78% identical or more, 79% identical or more, 80% identical or more, 81% identical or more, 82% identical or more, 83% identical or more, 84% identical or more, 85% identical or more, 86% identical or more, 87% identical or more, 88% identical or more, 89% identical or more, 90% identical or more, 91% identical or more, 92% identical or more, 93% identical or more, 94% identical or more, 95% identical or more, 96% identical or more, 97% identical or more, 98% identical or more, 99% identical or more, or100% identical to the transposase domain sequences provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350.

[0230] In some embodiments, the IS110 family transposase comprises a RuvC-like DEDD catalytic domain comprising an amino acid sequence that is 50% identical or more, 51% identical or more, 52% identical or more, 53% identical or more, 54% identical or more, 55% identical or more, 56% identical or more, 57% identical or more, 58% identical or more,59% identical or more, 60% identical or more, 61% identical or more, 62% identical or more,63% identical or more, 64% identical or more, 65% identical or more, 66% identical or more,67% identical or more, 68% identical or more, 69% identical or more, 70% identical or more,71% identical or more, 72% identical or more, 73% identical or more, 74% identical or more,75% identical or more, 76% identical or more, 77% identical or more, 78% identical or more,79% identical or more, 80% identical or more, 81% identical or more, 82% identical or more,83% identical or more, 84% identical or more, 85% identical or more, 86% identical or more,87% identical or more, 88% identical or more, 89% identical or more, 90% identical or more,91% identical or more, 92% identical or more, 93% identical or more, 94% identical or more,95% identical or more, 96% identical or more, 97% identical or more, 98% identical or more,99% identical or more, or 100% identical to the RuvC-like DEDD catalytic domain sequences provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524- 20350 and further comprises a transposase domain comprising an amino acid sequence that is 50% identical or more, 51% identical or more, 52% identical or more, 53% identical or more,54% identical or more, 55% identical or more, 56% identical or more, 57% identical or more,58% identical or more, 59% identical or more, 60% identical or more, 61% identical or more,62% identical or more, 63% identical or more, 64% identical or more, 65% identical or more,66% identical or more, 67% identical or more, 68% identical or more, 69% identical or more,70% identical or more, 71% identical or more, 72% identical or more, 73% identical or more,74% identical or more, 75% identical or more, 76% identical or more, 77% identical or more,78% identical or more, 79% identical or more, 80% identical or more, 81% identical or more,82% identical or more, 83% identical or more, 84% identical or more, 85% identical or more,86% identical or more, 87% identical or more, 88% identical or more, 89% identical or more,90% identical or more, 91% identical or more, 92% identical or more, 93% identical or more,94% identical or more, 95% identical or more, 96% identical or more, 97% identical or more,98% identical or more, 99% identical or more, or 100% identical to the transposase domain sequences provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524- 20350. In some embodiments, the amino acid sequence forms a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621 if the template modeling score (TM-score) for the RuvC-like DEDD catalytic domain is 0.5 or higher. In someembodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621 if the template modeling score (TM-score) is 0.5 or higher, 0.6 or higher, 0.7 or higher, 0.8 or higher, or 0.9 or higher. In some embodiments the RuvC-like DEDD catalytic domain sequence provided is the RuvC- like DEDD catalytic domain of “protein_IS621” of Figure 16 (SEQ ID NO: 10176). In some embodiments the RuvC-like DEDD catalytic domain sequence provided is the RuvC-like DEDD catalytic domain of “protein_IS1621_23122” of Figure 16 (SEQ ID NO: 10181). In some embodiments, the amino acid sequence forms a similar tertiary structure to the transposase domain of IS621. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 if the template modeling score (TM-score) for the transposase domain is 0.5 or higher. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 if the template modeling score (TM-score) is 0.5 or higher, 0.6 or higher, 0.7 or higher, 0.8 or higher, or 0.9 or higher. In some embodiments the transposase domain sequence provided is the transposase domain of “protein_IS621” of Figure 16 (SEQ ID NO: 10176). In some embodiments the transposase domain sequence provided is the transposase domain of “protein_IS1621_23122” of Figure 16 (SEQ ID NOT0181).

[0231] In some embodiments, the IS110 family transposase comprises a transposase domain comprising a motif G-x(28)-G-x(7)-SG-x(10)-G, G-x(29)-G-x(7)-SG-x(l 1)-G, IPGIG, IPGVG, AGLAP, LGLVP, [GSRAVECPYQNHFDKIT]-x(26,39)- [GKARCTNSEQHMD]-x(7,8)-[STRP][GASDNQCVER]-x(8,18)- [GACVYTSKRFNHIQDE], [GASFCRTNQKYDEV]-x(25,3 l)-[GRALKMSQCED]-x(7)- [ STRAP] [GN ADS VOTER] -x( 10,17)- [GSVRCAYKHQTFINED], [IVLAMFQTERCHSKYPWNG] [PKDTF YERVSHQ ANGILCMW] [GRASTCEPKVYQNH FDI]-x(0,l)-[IVAFLMCSWTYGKHNPRE]-x(0,2)-[GSDANKQERTMHCP], [IVLHAMFQTCYEPRSGW] [PRKES AD YTHVNGQICWLMF] [GSF AQC YTRKNWLDEV H][VIFLCMYATPSGW]-x(0,2)-[GADSQNKRTEPWHC],[ AVLNC STIGMFYWDHEPQR] [GRTKNACSEHQMD] [L VTIMAYSFCQRWKNPGH] [A CDTSVNERHFWIYQGKLMP][PVLAWISTCGNMR], or [LAVIFTCSWMGYPRQNKE] [GRLAMCKEQDS] [LIVTCMKS AFRPQG] [VTAICRNDGS QMHLEPYKF][PGKASVIDRLTENQCW] with motifs in common Prosite format, where xis any amino acid and x(n) represents n number of any amino acid and x(n,m) represents n to m number of any amino acids. In some embodiments, the transposase domain comprises a motif G-x(28)-G-x(7)-SG-x(10)-G, G-x(29)-G-x(7)-SG-x(l l)-G, IPGIG, IPGVG, AGLAP, or LGLVP. In some embodiments, the amino acid sequence forms a similar tertiary structure to the transposase domain of IS621. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 if the template modeling score (TM-score) for the transposase domain is 0.5 or higher. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 if the template modeling score (TM-score) is 0.5 or higher, 0.6 or higher, 0.7 or higher, 0.8 or higher, or 0.9 or higher.

[0232] In some embodiments, the IS110 family transposase comprises a transposase domain comprising an amino acid sequence that is 50% identical or more, 51% identical or more, 52% identical or more, 53% identical or more, 54% identical or more, 55% identical or more, 56% identical or more, 57% identical or more, 58% identical or more, 59% identical or more, 60% identical or more, 61% identical or more, 62% identical or more, 63% identical or more, 64% identical or more, 65% identical or more, 66% identical or more, 67% identical or more, 68% identical or more, 69% identical or more, 70% identical or more, 71% identical or more, 72% identical or more, 73% identical or more, 74% identical or more, 75% identical or more, 76% identical or more, 77% identical or more, 78% identical or more, 79% identical or more, 80% identical or more, 81% identical or more, 82% identical or more, 83% identical or more, 84% identical or more, 85% identical or more, 86% identical or more, 87% identical or more, 88% identical or more, 89% identical or more, 90% identical or more, 91% identical or more, 92% identical or more, 93% identical or more, 94% identical or more, 95% identical or more, 96% identical or more, 97% identical or more, 98% identical or more, 99% identical or more, or 100% identical to a transposase domain sequence of Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350 and comprising any of the motifs in the preceding paragraph. In some embodiments, the amino acid sequence forms a similar tertiary structure to the transposase domain of IS621. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 if the template modeling score (TM-score) forthe transposase domain is 0.5 or higher. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 if the template modeling score (TM-score) is 0.5 or higher, 0.6 or higher, 0.7 or higher, 0.8 or higher, or 0.9 or higher.

[0233] C.3. Linker Domain

[0234] The IS110 family transposases described herein comprise a linker domain between the RuvC-like DEDD catalytic domain and the transposase domain. In some embodiments, the linker domain comprised a coiled-coil.

[0235] In some embodiments, the IS110 family transposase comprises a linker domain comprising an amino acid sequence that is 50% identical or more, 51% identical or more, 52% identical or more, 53% identical or more, 54% identical or more, 55% identical or more, 56% identical or more, 57% identical or more, 58% identical or more, 59% identical or more, 60% identical or more, 61% identical or more, 62% identical or more, 63% identical or more, 64% identical or more, 65% identical or more, 66% identical or more, 67% identical or more, 68% identical or more, 69% identical or more, 70% identical or more, 71% identical or more, 72% identical or more, 73% identical or more, 74% identical or more, 75% identical or more, 76% identical or more, 77% identical or more, 78% identical or more, 79% identical or more, 80% identical or more, 81% identical or more, 82% identical or more, 83% identical or more, 84% identical or more, 85% identical or more, 86% identical or more, 87% identical or more, 88% identical or more, 89% identical or more, 90% identical or more, 91% identical or more, 92% identical or more, 93% identical or more, 94% identical or more, 95% identical or more, 96% identical or more, 97% identical or more, 98% identical or more, 99% identical or more, or 100% identical to the linker domain sequences provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350.

[0236] C.4. IS110 Transposases

[0237] In certain aspects, the invention provides an IS 110 family transposase comprising an N-terminal RuvC-like DEDD catalytic domain comprising a RuvC-like DEDD catalytic domain sequence provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350, a linker domain comprising a coiled-coil, and a C-terminal transposase domain comprising a transposase domain sequences provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350. In some embodiments, the linker domain comprises a linker domain sequence provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350. In some embodiments, the IS110 family transposase comprises“protein_IS621” of Figure 16 (SEQ ID NO: 10176). In some embodiments, the IS110 family transposase comprises “protein_IS1621_23122” of Figure 16 (SEQ ID NO: 10181).

[0238] In certain aspects, the invention provides an IS 110 family transposase consisting of an N-terminal RuvC-like DEDD catalytic domain consisting of a RuvC-like DEDD catalytic domain sequences provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350, a linker domain comprising a coiled-coil, and a C-terminal transposase domain consisting of a transposase domain sequences provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350. In some embodiments, the linker domain consists of a linker domain sequence provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350. In some embodiments, the IS 110 family transposase consists of “protein_IS621” of Figure 16 (SEQ ID NO: 10176). In some embodiments, the IS110 family transposase consists of “protein_IS1621_23122” of Figure 16 (SEQ ID NO: 10181).

[0239] In certain aspects, the invention provides an IS 110 family transposase comprising, in the N-terminal to C-terminal direction, a RuvC-like DEDD catalytic domain comprising any of the motifs or sequences described herein and 50% identical or more, 51% identical or more, 52% identical or more, 53% identical or more, 54% identical or more, 55% identical or more, 56% identical or more, 57% identical or more, 58% identical or more, 59% identical or more, 60% identical or more, 61% identical or more, 62% identical or more, 63% identical or more, 64% identical or more, 65% identical or more, 66% identical or more, 67% identical or more, 68% identical or more, 69% identical or more, 70% identical or more, 71% identical or more, 72% identical or more, 73% identical or more, 74% identical or more, 75% identical or more, 76% identical or more, 77% identical or more, 78% identical or more, 79% identical or more, 80% identical or more, 81% identical or more, 82% identical or more, 83% identical or more, 84% identical or more, 85% identical or more, 86% identical or more, 87% identical or more, 88% identical or more, 89% identical or more, 90% identical or more, 91% identical or more, 92% identical or more, 93% identical or more, 94% identical or more, 95% identical or more, 96% identical or more, 97% identical or more, 98% identical or more, 99% identical or more, or 100% identical to the RuvC-like DEDD catalytic domain sequences provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350, a linker domain comprising a coiled-coil, and a transposase domain comprising any of the motifs or sequences described herein and 50% identical or more, 51% identical or more, 52% identical or more, 53% identical or more, 54% identical or more, 55% identical or more, 56% identicalor more, 57% identical or more, 58% identical or more, 59% identical or more, 60% identical or more, 61% identical or more, 62% identical or more, 63% identical or more, 64% identical or more, 65% identical or more, 66% identical or more, 67% identical or more, 68% identical or more, 69% identical or more, 70% identical or more, 71% identical or more, 72% identical or more, 73% identical or more, 74% identical or more, 75% identical or more, 76% identical or more, 77% identical or more, 78% identical or more, 79% identical or more, 80% identical or more, 81% identical or more, 82% identical or more, 83% identical or more, 84% identical or more, 85% identical or more, 86% identical or more, 87% identical or more, 88% identical or more, 89% identical or more, 90% identical or more, 91% identical or more, 92% identical or more, 93% identical or more, 94% identical or more, 95% identical or more, 96% identical or more, 97% identical or more, 98% identical or more, 99% identical or more, or 100% identical to the transposase domain sequences provided in Figure 16 (SEQ ID NOS: 10176- 10523) or SEQ ID NOS: 10524-20350. In some embodiments, the linker domain comprises an amino acid sequence 50% identical or more, 51% identical or more, 52% identical or more, 53% identical or more, 54% identical or more, 55% identical or more, 56% identical or more, 57% identical or more, 58% identical or more, 59% identical or more, 60% identical or more, 61% identical or more, 62% identical or more, 63% identical or more, 64% identical or more, 65% identical or more, 66% identical or more, 67% identical or more, 68% identical or more, 69% identical or more, 70% identical or more, 71% identical or more, 72% identical or more, 73% identical or more, 74% identical or more, 75% identical or more, 76% identical or more, 77% identical or more, 78% identical or more, 79% identical or more, 80% identical or more, 81% identical or more, 82% identical or more, 83% identical or more, 84% identical or more, 85% identical or more, 86% identical or more, 87% identical or more, 88% identical or more, 89% identical or more, 90% identical or more, 91% identical or more, 92% identical or more, 93% identical or more, 94% identical or more, 95% identical or more, 96% identical or more, 97% identical or more, 98% identical or more, 99% identical or more, or 100% identical to the linker domain sequences provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350.

[0240] In certain aspects, described herein is an IS110 family transposase amino acid sequence further comprising a nuclear localization signal (NLS). In some embodiments the NLS is encoded at the N-terminus of the IS110 family transposase amino acid sequence. In some embodiments the NLS is encoded at the C-terminus of the IS110 family transposase amino acid sequence. In some embodiments, the IS110 family transposase amino acidsequence further comprises more than one NLS. In some embodiments, the IS110 family transposase amino acid sequence further comprises two NLSs. In some embodiments, the IS110 family transposase amino acid sequence further comprises three NLSs. In some embodiments, the IS110 family transposase amino acid sequence further comprises more than three NLSs. In some embodiments, the IS110 family transposase amino acid sequence further comprises three NLSs at the N-terminus of the IS110 family transposase amino acid sequence. In some embodiments, the IS110 family transposase amino acid sequence further comprises three NLSs at the C-terminus of the IS110 family transposase amino acid sequence. Any NLS known in that art can be used. In some embodiments, the NLS is an SV40 NLS.

[0241] The observed IS110 transposition reaction resembles that of conservative sitespecific recombinases (CSSRs). Two DNA sequences are specifically recognized by the transposition machinery. The bridgeRNA recognizes both the donor site sequence and target site sequence; the transposase binds the bridgeRNA and, in some embodiments, STIR subsequences of the donor site. Excision of an IS 110 transposon element is scarless, the original target sequence is reconstituted to its original sequence pre-insertion.

[0242] In some embodiments, the present invention could be supplemented with additional IS110 orthologs with improved efficiency and / or specificity using an extended bioinformatic search. IS110 systems with improved efficiency, such as in mammalian cells, could be identified by prioritizing candidates that occur in organisms that naturally grow at physiological temperatures. For example, candidates that naturally reside within human gut microbes could be prioritized for experimental characterization in human cells. Alternatively, IS110 systems that are found in the genomes of extremophile organisms could be prioritized for experimental characterization due to their potential utility as in vitro, temperature- responsive molecular tools. In another embodiment, IS110 systems with high specificity could be prioritized based on predicted properties of their bridgeRNA sequences, which are further described below (see Section D).

[0243] The present invention also contemplates the use of the IS110 system with cofactors and / or additional domains that may impart new functions for the IS110 system, increase the specificity of the IS110 system, and / or increase the efficiency of the IS110 system. These may exist within IS110 elements and / or in the vicinity of IS elements.

[0244] In some embodiments, the present invention contemplates an approach to engineering IS110 systems. Scientists have demonstrated that it is possible to infer theidentity of an ancestral protein using phylogenetic analysis (Alonso-Lerma, B., Jabalera, Y., Samperio, S. et al., Evolution of CRISPR-associated endonucleases as inferred from resurrected proteins., Nat Microbiol 8, 77-90 (2023)). These inferred ancestral proteins may not exist in nature in the present day, but they represent a common ancestor of an existing clade of proteins. These ancestral proteins could then be experimentally synthesized and tested to determine if they have favorable properties. In one embodiment, a clade of IS 110s with particular long or specific bridgeRNA target and / or donor binding loops could be analyzed using this method to identify an ancestral protein that is capable of binding to a bridgeRNA molecule with even more favorable properties.

[0245] In one embodiment, the present invention could be supplemented with additional IS110 orthologs using an extended bioinformatic search. Some IS110 transposases may have fused with other domains, imparting new biochemical properties to the IS110 system. In some embodiments, IS110 transposases with unusually long amino acid sequences could be investigated as potential candidates for these domain fusions. In some embodiments, protein domain collections, such as Pfam and InterPro, could be used to search across all proteins that are known to contain an IS110 RuvC-like DEDD catalytic domain, an IS 110 transposase domain, or both. These IS 110s could then be synthesized and experimentally characterized to identify any favorable properties that they may have. In another embodiment, IS110 transposases may not be fused to an additional protein cofactor as a single amino acid sequence, but they may associate with them in a multi-protein complex. These complexes could be identified by searching for proteins that occur near an IS 110 CDS with high frequency in natural genomes. These additional cofactors could be synthesized and experimentally characterized to identify any favorable properties that they may have.

[0246] In some embodiments, the present invention contemplates an approach to engineering IS110 systems wherein the transposase domain is replaced or fused with a integrase, a nucleobase deaminase, a reverse transcriptase, a recombinase, an integrase, a topoisomerase, a retrotransposon, phosphatase, polymerase, a ligase, a helitron, a helicase, a methylase, a demethylase, a translation activator, a translation repressor, a transcription activator, a transcription repressor, a transcription release factor, a chromatin modifier, a histone modifier, an acetylase, a deacetylase, a reverse transcriptase, or a nuclease. Fusions with the transposase domain can comprise simply fusing one of the aforementioned effector domain-comprising enzymes or domain(s) thereof with either the N- or C- terminus of the transposase. In some embodiments, the IS110 transposase may also be modified such that thecatalytic activity of the RuvC or transposase domain is crippled. In some embodiments, as an alternative to direct fusions, co-delivery of one or more of these effector domains along with, for example, an inactive IS110 transposase domain.

[0247] In some embodiments, the present invention contemplates an approach to engineering IS110 systems wherein the IS110 transposase comprises one or more amino acid mutations as compared to a wild type, whereby the mutations increase binding and / or interaction with a target site sequence, donor site sequence, and / or bridgeRNA and / or increase IS110 transposase activity

[0248] In some embodiments, the IS110 transposases may be engineered in such a way so as to reduce, limit, or eliminate the reverse reactions of insertion, excisive recombination, or inversion, in accordance with the observation that IS110 transposons excise themselves using an inefficient mechanism.

[0249] In certain aspects, described herein are nucleic acids encoding any of the IS110 family transposase amino acid sequences provided herein. In any of the embodiments described herein, a nucleotide sequence encoding the IS110 transposases or portions thereof can be codon-optimized. This type of optimization is known in the art and entails the mutation of foreign-derived DNA to mimic the codon preferences of the intended host organism or cell while encoding the same protein. Thus, the codons are changed, but the encoded protein remains unchanged. For example, if the intended target cell was a human cell, a human codon-optimized IS110 transposases or portions thereof would be a suitable IS110 transposase. As another non-limiting example, if the intended host cell were a mouse cell, then a mouse codon-optimized IS110 transposases or portions thereof would be a suitable IS110 transposase. While codon optimization is not required, it is acceptable and may be preferable in certain cases.

[0250] D. BridgeRNA

[0251] It was bioinformatically and experimentally determined that IS110 transposases bind an RNA, referred to herein as bridgeRNA or bRNA, expressed from the IS110 element non-coding ends and which directs donor and target site specificity. See Examples 1 and 3-5.

[0252] As discussed above, concatenation of the RE and LE may form a promoter which promotes expression of the bridgeRNA. The bridgeRNA forms an RNA-protein complex with the transposase and recognizes via base-pairing the target site or the donor site, the target site and donor site or two donor sites to mediate transposition. In someembodiments, the bridgeRNA can be recombinantly expressed using any suitable promoter. In some embodiments, portions of the bridgeRNA can be recombinantly expressed as two separate molecules using one promoter or multiple promoters. In some embodiments, the bridgeRNA can be recombinantly expressed using a promoter designed from the concatenation sequence of the RE and LE as described supra. In some embodiments, the bridgeRNA can be recombinantly expressed using a type III pol III promoter, which is known in the art to be suitable for small RNA expression, such as for guide RNAs in CRISPR-Cas systems. In some embodiments, the promoter is a sigma70 promoter box.

[0253] The bridgeRNA for IS110 transposases of the IS110 group is typically encoded within the LE, although in some embodiments the bridgeRNA can be at least partially encoded within RE also. In some embodiments, the bridgeRNA can be encoded within the LE and at least partially encoded within a 5' portion of the CDS sequence encoding the IS110 transposase. In some embodiments, the bridgeRNA for IS110 transposases of the IS110 group has donor and target binding loops to recognize the donor site sequence and the target site sequence via base-pairing to mediate transposition. In some embodiments, the bridgeRNA for IS110 transposases of the IS110 group has two or more binding loops that recognize two or more sequences via base-pairing to mediate transposition. If desired, one or more of the binding loops can be reprogrammed to recognize any sequence of interest.

[0254] The bridgeRNA for IS110 transposases of the IS1111 group is typically encoded within the RE. In some embodiments, the bridgeRNA can be encoded within the RE and at least partially encoded within a 3' portion of the CDS sequence encoding the IS110 transposase (e.g., an IS1111 group IS110 transposase). In some embodiments, the bridgeRNA for IS110 transposases of the IS1111 group have donor and target binding loops to recognize the donor site sequence and the target site sequence via base-pairing to mediate transposition. In some embodiments, the donor and target binding sequences of the bridgeRNA for IS110 transposases of the IS1111 group can be present on a single loop of the bridgeRNA that recognizes the donor site sequence and the target site sequence via basepairing to mediate transposition. In some embodiments, the donor binding sequences of the bridgeRNA for IS110 transposases of the IS1111 group can be present on a multi -branched loop of the bridgeRNA that recognizes the donor site sequence via base-pairing to mediate transposition. In some embodiments, the bridgeRNA for IS110 transposases of the IS1111 group has two or more binding loops that recognize two or more sequences via base-pairing to mediate transposition. If desired, one or more of the binding loops can be reprogrammedto recognize any sequence of interest.

[0255] In some embodiments, bridgeRNAs for use with IS110 transposases can comprise one or more donor binding loops to recognize two identical or different donor site sequences via base-pairing to mediate transposition. In some embodiments, the bridgeRNA does not comprise a target binding loop. In some embodiments, the bridgeRNA further comprises a target binding loop.

[0256] The bridgeRNA refers to a single-stranded RNA molecule that undergoes intramolecular base pairing to form one or more stem-loop structures (which includes stem structures with external loops, bulge loops, multi-branched loops, hairpin loops and / or internal loops, see e.g., Fig 1 of Sato, K., Akiyama, M. & Sakakibara, Y. RNA secondary structure prediction using deep learning with thermodynamic integration. Nat Commun 12, 941 (2021), doi. org / 10.1038 / s41467-021-21194-4, the content of which is hereby incorporated by reference in its entirety). Typically, a stem is formed by two portions of the RNA molecule which are complementary when read in opposite directions and base pair to form a stem with intervening unpaired nucleotides between the two complementary portions forming a loop. A stem sequence does not need to be 100% complementary and can comprise mismatches, bulges, or internal loops.

[0257] In some embodiments, the bridgeRNA comprises a RNA molecule that comprises at least one stem-loop structure and further comprises one or more loops comprising a first nucleotide sequence that is complementary to a first target site sequence of a target DNA, a second nucleotide sequence that is complementary to a second target site sequence which is on the opposite strand of the target DNA to the first target site sequence and a handshake guide nucleotide sequence that is complementary to a donor site sequence, a third nucleotide sequence that is complementary to a first donor site sequence of a donor DNA, and a fourth nucleotide sequence that is complementary to a second donor site sequence which is on the opposite strand of the donor DNA to the first donor site sequence, and a handshake guide nucleotide sequence that is complementary to a target site sequence. In some embodiments, the bridgeRNA comprises a first internal loop comprising a first nucleotide sequence that is complementary to a first target site sequence of a target DNA, a second nucleotide sequence that is complementary to a second target site sequence which is on the opposite strand of the target DNA to the first target site sequence, and a handshake guide nucleotide sequence that is complementary to a donor site sequence, and a second internal loop comprising a third nucleotide sequence that is complementary to a first donorsite sequence of a donor DNA, and a fourth nucleotide sequence that is complementary to a second donor site sequence which is on the opposite strand of the donor DNA to the first donor site sequence and a handshake guide nucleotide sequence that is complementary to a target site sequence. In some embodiments, the bridgeRNA comprises an internal loop comprising a first nucleotide sequence that is complementary to a first target site sequence of a target DNA, a second nucleotide sequence that is complementary to a second target site sequence which is on the opposite strand of the target DNA to the first target site sequence and a handshake guide nucleotide sequence that is complementary to a donor site sequence, and a multi-branched loop comprising a third nucleotide sequence that is complementary to a first donor site sequence of a donor DNA, and a fourth nucleotide sequence that is complementary to a second donor site sequence which is on the opposite strand of the donor DNA to the first donor site sequence and a handshake guide nucleotide sequence that is complementary to a target site sequence. In some embodiments, said bridgeRNA binds an IS110 group transposase. In some embodiments, said bridgeRNA binds an IS1111 group transposase. In some embodiments, said bridgeRNA binds the transposase of IS621, ISPal 1, IsPa29, ISMmgl, ISPfll, ISMae40, ISStma6, ISAzs32, ISMex9, ISCARN28, ISAarl6, ISCps7, ISPpu9, ISRel9, ISEsa2, ISMma5, IS900, or ISHne5. In some embodiments, said bridgeRNA binds the transposase of IS621. In some embodiments, a loop comprising the first and second nucleotide sequences that are complementary to the target site sequence of a target DNA may form a stem or partial stem structure when not bound by a transposase and form a single stranded structure when bound to a transposase. In some embodiments, a loop comprising the third and fourth nucleotide sequences that are complementary to the donor site sequence of a donor DNA may form a stem or partial stem structure when not bound by a transposase and form a single stranded structure when bound to a transposase.

[0258] With reference to the above bridgeRNA, in some embodiments, the first and second nucleotide sequences are fully complementary to their respective target site sequences of the target DNA. In some embodiments, the first and / or second nucleotide sequences are partially complementary to their respective target site sequences of the target DNA, i.e., there is non-canonical base pairing, mismatch and / or non-contiguous tolerance. In some embodiments, there is a single mismatch in the first or second nucleotide sequence. In some embodiments, there are two, three or four mismatches which can be in the first nucleotide sequence, in the second nucleotide sequence, or spread across the first and second nucleotide sequences. In some embodiments, the two, three or four mismatches are contiguous. Insome embodiments, the two, three or four mismatches are non-contiguous. In some embodiments, there is a single non-canonical base pair in the first or second nucleotide sequence. In some embodiments, there are two, three or four non-canonical base pairs which can be in the first nucleotide sequence, in the second nucleotide sequence, or spread across the first and second nucleotide sequences. In some embodiments, the two, three or four non- canonical base pairs are contiguous. In some embodiments, the two, three or four non- canonical base pairs are non-contiguous. In some embodiments, the third and fourth nucleotide sequences are fully complementary to their respective donor site sequences of the donor DNA. In some embodiments, the third and / or fourth nucleotide sequences are partially complementary to their respective donor site sequences of the donor DNA, i.e., there is non- canonical base pairing, mismatch and / or non-contiguous tolerance. In some embodiments, there is a single mismatch in the third or fourth nucleotide sequence. In some embodiments, there are two mismatches which can be in the third nucleotide sequence, in the fourth nucleotide sequence, or spread across the third and fourth nucleotide sequences. In some embodiments, the two mismatches are contiguous. In some embodiments, the two mismatches are non-contiguous. In some embodiments, there is a single non-canonical base pair in the third or fourth nucleotide sequence. In some embodiments, there are two non- canonical base pairs which can be in the third nucleotide sequence, in the fourth nucleotide sequence, or spread across the third and fourth nucleotide sequences. In some embodiments, the two non-canonical base pairs are contiguous. In some embodiments, the two non- canonical base pairs are non-contiguous. In some embodiments, the non-canonical base pairing is non-Watson-Crick base pairing (i.e., is not a G-C base pair or A-T / U base pair). For example, see Pacesa M, Lin CH, Clery A, et al., Structural basis for Cas9 off-target activity, Cell, 2022, 185(22):4067-4081.e21, the contents of which is hereby incorporated by reference in its entirety). In some embodiments, the non-canonical base pairing is wobble base pairing. In some embodiments, the non-canonical base pairing is Hoogsteen base pairing. In some embodiments, the non-canonical base pairing comprises a rG-dT base pair, a rU-dG base pair, a rA-dC base pair, a rC-dA base pair, a rA-dG base pair, or a rG-dG base pair.

[0259] With reference to the above HSGs (e g., HSG-TBL and / or HSG-DBL), in some embodiments, the HSG nucleotide sequences are fully complementary to their associated anti-HSG sequences of the anti-HSG sequence site. In some embodiments, the HSG sequences are partially complementary to their associated anti-HSG sequences of theanti-HSG sequence site, i.e., there is non-canonical base pairing, mismatch and / or noncontiguous tolerance. In some embodiments, there is a single mismatch in the HSG nucleotide sequence. In some embodiments, there are two, three or four mismatches. In some embodiments, the two, three or four mismatches are contiguous. In some embodiments, the two, three or four mismatches are non-contiguous. In some embodiments, there is a single non-canonical base pair in the HSG nucleotide sequence. In some embodiments, there are two, three or four non-canonical base pairs. In some embodiments, the two, three or four non-canonical base pairs are contiguous. In some embodiments, the two, three or four non-canonical base pairs are non-contiguous. In some embodiments, the non-canonical base pairing is non-Watson-Crick base pairing (i.e., is not a G-C base pair or A-T / U base pair). For example, see Pacesa M, Lin CH, Clery A, et al., Structural basis for Cas9 off-target activity, Cell, 2022, 185(22):4067-4081.e21, the contents of which is hereby incorporated by reference in its entirety). In some embodiments, the non-canonical base pairing is wobble base pairing. In some embodiments, the non-canonical base pairing is Hoogsteen base pairing. In some embodiments, the non-canonical base pairing comprises a rG-dT base pair, a rU-dG base pair, a rA-dC base pair, a rC-dA base pair, a rA-dG base pair, or a rG-dG base pair.

[0260] In some embodiments, the bridgeRNA comprises a RNA molecule that comprises at least one stem-loop structure and further comprises one or more loops comprising a first nucleotide sequence that is complementary to a first target site sequence of a target DNA, a second nucleotide sequence that is complementary to a second target site sequence which is on the opposite strand of the target DNA to the first target site sequence and a handshake guide nucleotide sequence that is complementary to a donor site sequence. In some embodiments, the bridgeRNA further comprises a third nucleotide sequence that is complementary to a first donor site sequence of a donor DNA, and a fourth nucleotide sequence that is complementary to a second donor site sequence which is on the opposite strand of the donor DNA to the first donor site sequence and a handshake guide nucleotide sequence that is complementary to a target site sequence. In some embodiments, the bridgeRNA comprises a first internal loop comprising a first nucleotide sequence that is complementary to a first target site sequence of a target DNA, a second nucleotide sequence that is complementary to a second target site sequence which is on the opposite strand of the target DNA to the first target site sequence and a handshake guide nucleotide sequence that is complementary to a donor site sequence. In some embodiments, the bridgeRNA comprises asecond internal loop comprising a third nucleotide sequence that is complementary to a first donor site sequence of a donor DNA, and a fourth nucleotide sequence that is complementary to a second donor site sequence which is on the opposite strand of the donor DNA to the first donor site sequence and a handshake guide nucleotide sequence that is complementary to a target site sequence. In some embodiments, the first internal loop (e.g. a multi-branched loop) of the bridgeRNA comprises a third nucleotide sequence that is complementary to a first donor site sequence of a donor DNA, and a fourth nucleotide sequence that is complementary to a second donor site sequence which is on the opposite strand of the donor DNA to the first donor site sequence and a handshake guide nucleotide sequence that is complementary to a target site sequence. In some embodiments, said bridgeRNA binds an IS1111 group transposase. In some embodiments, said bridgeRNA binds the transposase of IS1111 229727. In some embodiments, a loop comprising the first and second nucleotide sequences that are complementary to the target site sequence of a target DNA may form a stem or partial stem structure when not bound by a transposase and form a single stranded structure when bound to a transposase. In some embodiments, a loop comprising the third and fourth nucleotide sequences that are complementary to the donor site sequence of a donor DNA may form a stem or partial stem structure when not bound by a transposase and form a single stranded structure when bound to a transposase.

[0261] With reference to the above bridgeRNA, in some embodiments, the first and second nucleotide sequences are fully complementary to their respective target site sequences of the target DNA. In some embodiments, the first and / or second nucleotide sequences are partially complementary to their respective target site sequences of the target DNA, i.e., there is non-canonical base pairing, mismatch and / or non-contiguous tolerance. In some embodiments, there is a single mismatch in the first or second nucleotide sequence. In some embodiments, there are two, three or four mismatches which can be in the first nucleotide sequence, in the second nucleotide sequence, or spread across the first and second nucleotide sequences. In some embodiments, the two, three or four mismatches are contiguous. In some embodiments, the two, three or four mismatches are non-contiguous. In some embodiments, there is a single non-canonical base pair in the first or second nucleotide sequence. In some embodiments, there are two, three or four non-canonical base pairs which can be in the first nucleotide sequence, in the second nucleotide sequence, or spread across the first and second nucleotide sequences. In some embodiments, the two, three or four non- canonical base pairs are contiguous. In some embodiments, the two, three or four non-canonical base pairs are non-contiguous. In some embodiments, the third and fourth nucleotide sequences are fully complementary to their respective donor site sequences of the donor DNA. In some embodiments, the third and / or fourth nucleotide sequences are partially complementary to their respective donor site sequences of the donor DNA, i.e., there is non- canonical base pairing, mismatch and / or non-contiguous tolerance. In some embodiments, there is a single mismatch in the third or fourth nucleotide sequence. In some embodiments, there are two mismatches which can be in the third nucleotide sequence, in the fourth nucleotide sequence, or spread across the third and fourth nucleotide sequences. In some embodiments, the two mismatches are contiguous. In some embodiments, the two mismatches are non-contiguous. In some embodiments, there is a single non-canonical base pair in the third or fourth nucleotide sequence. In some embodiments, there are two non- canonical base pairs which can be in the third nucleotide sequence, in the fourth nucleotide sequence, or spread across the third and fourth nucleotide sequences. In some embodiments, the two non-canonical base pairs are contiguous. In some embodiments, the two non- canonical base pairs are non-contiguous. In some embodiments, the non-canonical base pairing is non-Watson-Crick base pairing (i.e., is not a G-C base pair or A-T / U base pair). In some embodiments, the non-canonical base pairing is wobble base pairing. In some embodiments, the non-canonical base pairing is Hoogsteen base pairing. In some embodiments, the non-canonical base pairing comprises a rG-dT base pair, a rU-dG base pair, a rA-dC base pair, a rC-dA base pair, a rA-dG base pair, or a rG-dG base pair.

[0262] With reference to the above HSGs (e g., HSG-TBL and / or HSG-DBL), in some embodiments, the HSG nucleotide sequences are fully complementary to their associated anti-HSG sequences of the anti-HSG sequence site. In some embodiments, the HSG sequences are partially complementary to their associated anti-HSG sequences of the anti-HSG sequence site, i.e., there is non-canonical base pairing, mismatch and / or noncontiguous tolerance. In some embodiments, there is a single mismatch in the HSG nucleotide sequence. In some embodiments, there are two, three or four mismatches. In some embodiments, the two, three or four mismatches are contiguous. In some embodiments, the two, three or four mismatches are non-contiguous. In some embodiments, there is a single non-canonical base pair in the HSG nucleotide sequence. In some embodiments, there are two, three or four non-canonical base pairs. In some embodiments, the two, three or four non-canonical base pairs are contiguous. In some embodiments, the two, three or four non-canonical base pairs are non-contiguous. In some embodiments, thenon-canonical base pairing is non-Watson-Crick base pairing (i.e., is not a G-C base pair or A-T / U base pair). For example, see Pacesa M, Lin CH, Clery A, et al., Structural basis for Cas9 off-target activity, Cell, 2022, 185(22):4067-4081.e21, the contents of which is hereby incorporated by reference in its entirety). In some embodiments, the non-canonical base pairing is wobble base pairing. In some embodiments, the non-canonical base pairing is Hoogsteen base pairing. In some embodiments, the non-canonical base pairing comprises a rG-dT base pair, a rU-dG base pair, a rA-dC base pair, a rC-dA base pair, a rA-dG base pair, or a rG-dG base pair.

[0263] In some embodiments, the bridgeRNA comprises a RNA molecule that comprises at least two stem-loop structures, and with respect to the bridgeRNA having at least two stem-loop structures, these two structures are referred to as the “first stem-loop” and the “second stem-loop” where the first is 5' or upstream to the second. In some embodiments, the first stem-loop comprises a target binding loop and the second stem-loop comprises a donor binding loop. In some embodiments, the first stem loop comprises a donor binding loop and the second stem-loop comprises a target binding loop. Thus, with this nomenclature, there can be additional stem-loop structures upstream of the first stem-loop, between the first and second stem-loops, and / or after the second-stem loop. In some embodiments, the bridgeRNA comprises at least a first internal loop referred to as a target binding loop and a second internal loop referred to as a donor binding loop. In some embodiments, the bridgeRNA comprises a RNA molecule that comprises at least two stemloop structures as depicted in Figure 2D and for Cluster 2 in Figure 13. In Figure 2D and for Cluster 2 in Figure 13, and in some embodiments, the first stem-loop structure of the bridgeRNA comprises a first stem-loop (5-35 nt, 3-10 nt loop) comprising an internal loop (e.g., a target binding loop) (5-20 nt). In Figure 2D and for Cluster 2 in Figure 13, and in some embodiments, the second stem-loop structure of the bridgeRNA comprises a second stem-loop (5-35 nt, 3-10 nt loop) comprising an internal loop (e.g., a donor binding loop) (5- 20 nt). The stem of the second stem-loop structure can include additional loops and bubbles that are 1-10 nucleotides each. As shown in Figure 2D and in some embodiments, the bridgeRNA may further comprise an additional third stem-loop structure (5-15 nt stem, 3-10 nt loop) (accessory structure) 5' to the first stem-loop. In some embodiments, said bridgeRNA binds the transposase of IS621. In some embodiments, said bridgeRNA binds the transposase of ISMae40 or ISStma6. In some embodiments, the donor binding loop and / or target binding loop may form a stem or partial stem structure when not bound by atransposase and form a single stranded structure when bound to a transposase.

[0264] Thus, in some embodiments, the bridgeRNA comprises a nucleotide sequence comprising the following secondary structure: first stem-loop comprising an internal target binding loop - second stem-loop comprising an internal donor binding loop. In some embodiments, the bridgeRNA comprises additional stem-loop structures, bulges, and / or loops (see, e.g., Fig. 2D). In some embodiments, the bridgeRNA comprises a nucleotide sequence comprising 5'-[A]-[B]-[C]-[D]-[E]-[F]-[G]-[H]-[I]-[J]-[K]-[L]-[M]-[N]-3', wherein A is a first stem portion, B is a first side of an internal loop corresponding to a target binding loop, C is a second stem portion, D is a first loop portion, E is the reverse complement of C, F is a second side of the internal loop corresponding to the target binding loop, G is the reverse complement of A, H is a third stem portion, l is a first side of an internal loop corresponding to a donor binding loop, J is a fourth stem portion, K is a second loop portion, L is the reverse complement of J, M is a second side of the internal loop corresponding to the donor binding loop, and N is the reverse complement of H. In some embodiments, the reverse complement portions are not 100% complementary so that the stem structures may comprise one or more mismatches or bulges, or non-standard base-pairing may occur. In some embodiments, the first side of the internal loop corresponding to the target binding loop comprises a nucleotide sequence that is complementary to a first target site sequence of a target DNA (referred to as LTG (left target guide)) and the second side of the internal loop corresponding to the target bind loop comprises a second nucleotide sequence that is complementary to a second target site sequence which is on the opposite strand of the target DNA to the first target site sequence (referred to as RTG (right target guide)) and comprises a handshake guide nucleotide sequence that is complementary to a donor site sequence which is on the opposite strand of the donor DNA strand bound by the LDG. In some embodiments, the first side of the internal loop corresponding to the target bind loop comprises a nucleotide sequence that is complementary to a first target site sequence of a target DNA (referred to as LTG) and comprises a handshake guide nucleotide sequence that is complementary to a donor site sequence which is on the opposite strand of the donor DNA strand bound by the RDG and the second side of the internal loop corresponding to the target bind loop comprises a second nucleotide sequence that is complementary to a second target site sequence which is on the opposite strand of the target DNA to the first target site sequence (referred to as RTG). In some embodiments, the first side of the internal loop corresponding to the donor binding loop comprises a nucleotide sequence that is complementary to a first donor site sequence of adonor DNA (referred to as LDG (left donor guide)) and the second side of the internal loop corresponding to the donor bind loop comprises a second nucleotide sequence that is complementary to a second donor site sequence which is on the opposite strand of the donor DNA to the first target site sequence (referred to as RDG (right donor guide)) and comprises a handshake guide nucleotide sequence that is complementary to a target site sequence which is on the opposite strand of the target DNA strand bound by the LTG. In some embodiments, the first side of the internal loop corresponding to the donor binding loop comprises a nucleotide sequence that is complementary to a first donor site sequence of a donor DNA (referred to as LDG and comprises a handshake guide nucleotide sequence that is complementary to a target site sequence which is on the opposite strand of the target DNA strand bound by the RTG and the second side of the internal loop corresponding to the donor bind loop comprises a second nucleotide sequence that is complementary to a second donor site sequence which is on the opposite strand of the donor DNA to the first target site sequence (referred to as RDG). In some embodiments, the stem structures may comprise one or more mismatches or bulges. In some embodiments, at least two portions of nucleotide sequence N are not complementary to the nucleotide sequence of portion H, so that the stem structure formed by base pairing between portion H and N comprises two bulges. In some embodiments, additional nucleotides are present between any of the portions A to N. For example, one or more non-base-paired nucleotides are present between portions G and H. In another embodiment, one or more nucleotides are present between the 5' terminus and portion A. In another embodiment, one or more nucleotides are present between the 3' terminus and portion N. In some embodiments, said bridgeRNA binds the transposase of IS621. In some embodiments, said bridgeRNA binds the transposase of ISMae40 or ISStma6. In some embodiments, the donor binding loop and / or target binding loop may form a stem or partial stem structure when not bound by a transposase and form a single stranded structure when bound to a transposase.

[0265] In some embodiments, the bridgeRNA comprises a nucleotide sequence comprising the following secondary structure: third stem-loop - first stem-loop comprising an internal target binding loop - second stem-loop comprising an internal donor binding loop. In some embodiments, the bridgeRNA comprises additional stem-loop structures, bulges, and / or loops (see, e.g., Fig. 2D). In some embodiments, the bridgeRNA comprises a nucleotide sequence comprising 5'-[Z]-[X]-[Y]-[A]-[B]-[C]-[D]-[E]-[F]-[G]-[H]-[I]-[J]-[K]-[L]- [M]-[N]-3', wherein Z is a fifth stem portion, X is a third loop portion, Y is the reversecomplement of Z, A is a first stem portion, B is a first side of an internal loop corresponding to a target binding loop, C is a second stem portion, D is a first loop portion, E is the reverse complement of C, F is a second side of the internal loop corresponding to the target binding loop, G is the reverse complement of A, H is a third stem portion, I is a first side of an internal loop corresponding to a donor binding loop, J is a fourth stem portion, K is a second loop portion, L is the reverse complement of J, M is a second side of the internal loop corresponding to the donor binding loop, and N is the reverse complement of H. In some embodiments, the reverse complement portions are not 1...

Claims

What is claimed is:

1. A recombinant nucleic acid editing system comprising: a) an IS110 family transposase, or a nucleic acid comprising a sequence encoding the IS110 family transposase; and b) a nucleic acid comprising a sequence encoding a bridgeRNA, wherein the bridgeRNA comprises one or more handshake guide (HSG) sequences.

2. The recombinant nucleic acid editing system of claim 1, wherein the nucleic acid comprising a sequence encoding a bridgeRNA comprises a left end (LE) sequence of a transposon that encodes the IS110 family transposase.

3. The recombinant nucleic acid editing system of claim 1, wherein the nucleic acid comprising a sequence encoding a bridgeRNA comprises a right end (RE) sequence of a transposon that encodes the IS110 family transposase.

4. The recombinant nucleic acid editing system of claims 1-3, further comprising a nucleic acid comprising a RE sequence, an anti-HSG donor site sequence, and a LE sequence, or a RE sequence, an anti-HSG donor site sequence, a core sequence, and a LE sequence, of the IS110 element that encodes the IS110 family transposase.

5. The recombinant nucleic acid editing system of claims 1-3, further comprising a nucleic acid comprising a right flank (RF) sequence, an anti-HSG target site sequence, and a left flank (LF) sequence, or a RF sequence, an anti-HSG target site sequence, a core sequence, and a LF sequence, of the target site sequence for the IS110 family transposase.

6. The recombinant nucleic acid editing system of claim 4, wherein the nucleic acid comprising the RE sequence, the anti-HSG donor site sequence, and the LE sequence, or the RE sequence, the anti-HSG donor site sequence, the core sequence, and the LE sequence, further comprises a nucleic acid sequence for insertion into a target site sequence.

7. The recombinant nucleic acid editing system of claim 6, wherein the target site sequence comprises a RF sequence, an anti-HSG target site sequence, and a LF sequence, or a RF sequence, an anti-HSG target site sequence, a core sequence, and a LF sequence for the IS110 family transposase.

8. The recombinant nucleic acid editing system of claim 5, wherein the nucleic acid comprising the RF sequence, the anti-HSG target site sequence, and the LF sequence, or the RF sequence, the anti-HSG target site sequence, the core sequence, and the LF sequence, further comprises a nucleic acid sequence for insertion into a donor site sequence.

9. The recombinant nucleic acid editing system of claim 8, wherein the donor site sequence comprises a RE sequence, an anti-HSG donor site sequence, and a LE sequence, or a RE sequence, an anti-HSG donor site sequence, a core sequence, and a LE sequence, of the IS110 element that encodes the IS110 family transposase.

10. The recombinant nucleic acid editing system of claims 1-9, wherein the bridgeRNA comprises a nucleotide sequence at least 50% identical to a bridgeRNA sequence of SEQ ID NOS: 1-348 or SEQ ID NOS: 349-10175.

11. The recombinant nucleic acid editing system of claims 1-9, wherein said RE sequence, comprises a RE sequence of SEQ ID NOS: 1-348, 30354-30529, 349-10175 or 30530-40356, said LE sequence comprises a LE sequence of SEQ ID NOS: 1-348, 30354- 30529, 349-10175 or 30530-40356, and / or said core sequence comprises a core sequence of SEQ ID NOS: 1-348, 30354-30529, 349-10175 or 30530-40356.

12. A recombinant nucleic acid editing system comprising: a) an IS110 family transposase, or a nucleic acid comprising a sequence encoding the IS110 family transposase; and b) a bridgeRNA, or a nucleic acid comprising a sequence encoding the bridgeRNA, the bridgeRNA comprising at least one stem-loop structure and further comprising at least one internal loop comprising a first nucleotide sequence that is complementary to a first target site sequence of a target DNA, a second nucleotide sequence that is complementary to a second target site sequence which is on the opposite strand of the target DNA to the first target site sequence, and a handshake guide nucleotide sequence that is complementary to a donor site sequence, and wherein the bridgeRNA is capable of forming a complex with the IS110 family transposase.

13. The recombinant nucleic acid editing system of claim 12, wherein the bridgeRNA further comprises a third nucleotide sequence that is complementary to a first donor sitesequence of a donor DNA, a fourth nucleotide sequence that is complementary to a second donor site sequence which is on the opposite strand of the donor DNA to the first donor site sequence, and a handshake guide nucleotide sequence that is complementary to a target site sequence.

14. The recombinant nucleic acid editing system of claim 13, wherein the third nucleotide sequence, the fourth nucleotide sequence, and handshake guide nucleotide sequence that is complementary to a target site sequence are on a second internal loop.

15. The recombinant nucleic acid editing system of claims 1-14, wherein the IS110 family transposase comprises a RuvC-like DEDD catalytic domain and a transposase domain.

16. The recombinant nucleic acid editing system of claim 15, wherein the IS110 family transposase further comprises a linker domain between the RuvC-like DEDD catalytic domain and transposase domain.

17. The recombinant nucleic acid editing system of claim 16, wherein the linker domain comprises a coiled-coil linker domain.

18. The recombinant nucleic acid editing system of claims 15-17, wherein the RuvC- like DEDD catalytic domain comprises an amino acid sequence at least 50% identical to a RuvC-like DEDD catalytic domain sequence of SEQ ID NOs: 10176-10523 or 10524-20350.

19. The recombinant nucleic acid editing system of claims 15-17, wherein the IS110 family transposase comprises a RuvC-like DEDD catalytic domain that forms a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621.

20. The recombinant nucleic acid editing system of claim 19, wherein the IS110 family transposase RuvC-like DEDD catalytic domain comprises a tertiary structure similar to a tertiary structure of the RuvC-like DEDD catalytic domain of IS621 if the template modeling score (TM-score) for the RuvC-like DEDD catalytic domain of the IS110 family transposase is 0.5 or higher.

21. The recombinant nucleic acid editing system of claim 20, wherein the RuvC-like DEDD catalytic domain comprises an amino acid sequence at least 15% identical to a RuvC- like DEDD catalytic domain sequence of SEQ ID NOs: 10176-10523 or 10524-20350.

22. The recombinant nucleic acid editing system of claims 15-17, wherein the transposase domain comprises an amino acid sequence at least 50% identical to a transposase domain sequence of SEQ ID NOs: 10176-10523 or 10524-20350.

23. The recombinant nucleic acid editing system of claims 15-17, wherein the IS110 family transposase comprises a transposase domain that forms a similar tertiary structure to the transposase domain of IS621.

24. The recombinant nucleic acid editing system of claim 23, wherein the IS110 family transposase domain comprises a tertiary structure similar to a tertiary structure of the transposase domain of IS621 if the template modeling score (TM-score) for the transposase domain of the IS110 family transposase is 0.5 or higher.

25. The recombinant nucleic acid editing system of claim 24, wherein the transposase domain comprises an amino acid sequence at least 15% identical to a transposase domain sequence of SEQ ID NOs: 10176-10523 or 10524-20350.

26. The recombinant nucleic acid editing system of claims 15-17, wherein the IS110 family transposase comprises an amino acid sequence at least 50% identical to SEQ ID NOs: 10176-10523 or 10524-20350.

27. The recombinant nucleic acid editing system of claims 15-17, wherein the IS110 family transposase further comprises a tertiary structure similar to a tertiary structure of IS621.

28. The recombinant nucleic acid editing system of claim 27, wherein the IS110 family transposase comprises a tertiary structure similar to a tertiary structure of IS621 if the template modeling score (TM-score) for the transposase is 0.5 or higher.

29. The recombinant nucleic acid editing system of claim 28, wherein the transposase domain comprises an amino acid sequence at least 15% identical to a transposase domain sequence of SEQ ID NOs: 10176-10523 or 10524-20350.

30. The recombinant nucleic acid editing system of claims 1-29, wherein the IS110 family transposase is an IS 110 group transposase.

31. The recombinant nucleic acid editing system of claims 1-29, wherein the IS110 family transposase is an IS1111 group transposase.

32. The recombinant nucleic acid editing system of claim 31, wherein the IS1111 group transposase is IS1111 A or IS1111_229727.

33. The recombinant nucleic acid editing system of claim 30, wherein the IS110 group transposase is IS621, ISPal l, IsPa29, ISMmgl, ISPfll, ISMae40, ISStma6, ISAzs32, ISMex9, ISCARN28, IS Aar 16, ISCps7, ISPpu9, ISRel9, ISEsa2, ISMma5, IS900, or ISHne5.

34. The recombinant nucleic acid editing system of claim 30, wherein the IS110 group transposase comprises an amino acid sequence at least 50% identical to IS621 (SEQ ID NO: 10176).

35. The recombinant nucleic acid editing system of claims 12-34, wherein the bridgeRNA comprises at least two stem-loop structures comprising a first stem-loop and a second stem-loop where the first stem-loop is 5' to the second stem-loop and wherein the first stem-loop comprises a target binding loop and the second stem-loop comprises a donor binding loop.

36. The recombinant nucleic acid editing system of claim 35, wherein the bridgeRNA further comprises a third stem-loop structure 5' of the first stem-loop.

37. The recombinant nucleic acid editing system of claims 35-36, wherein the stem of the first stem-loop is 5 to 35 nucleotides long and the loop is 3-10 nucleotides long, the target binding loop is 5 to 20 nucleotides long, the stem of the second stem-loop is 5 to 35 nucleotides long and the loop is 3-10 nucleotides long, and the donor binding loop is 5 to 20 nucleotides long.

38. The recombinant nucleic acid editing system of claim 37, wherein the stem of the second stem-loop structure comprises 1 to 4 loops or bubbles that are each 1 to 10 nucleotides long.

39. The recombinant nucleic acid editing system of claims 12-34, wherein the bridgeRNA comprises a nucleotide sequence comprising any of the 5' to 3' sequencesprovided in Figure 19, wherein “n” represents any nucleotide, “R” represents an A or G nucleotide, and “Y” represents a C or U nucleotide.

40. The recombinant nucleic acid editing system of claim 39, wherein the bridgeRNA comprises a 5' to 3' secondary structure provided in the first row, second row, third row, or fourth row of secondary structure for said sequence provided in Figure 19, wherein matching parentheses “(“ and “)” indicate base-paired nucleotides, and indicate unpaired bases.

41. The recombinant nucleic acid editing system of claims 12-34, wherein the bridgeRNA comprises a stem-loop structure as depicted in Figure 2D, Figure 1 IB, or Figure 13.

42. The recombinant nucleic acid editing system of claims 12-41, wherein the target binding loop of the bridgeRNA comprises: a left-target guide (LTG) comprising, in the 5' to 3' direction, a nucleotide sequence complementary to first strand of a target site sequence wherein the 3' end of the LTG is complementary to at least one of the nucleotides of a core sequence on the first strand of the target site sequence; a right-target guide (RTG) comprising, in the 5' to 3' direction, a nucleotide sequence that is a reverse complement to an opposite strand of the first strand of the target site sequence wherein the 3' end of the RTG is reverse complementary to at least one of the nucleotides of the core sequence on the opposite strand of the first strand of the target site sequence; and a handshake guide (HSG-TBL) comprising, in the 5' to 3' direction, a nucleotide sequence that is partially or fully, complementary to the opposite strand of the first strand of the donor site sequence (anti-HSG-donor site sequence) and wherein optionally the HSG-TBL is located 3' to the 3' end of the RTG or wherein the HSG-TBL is located 3' of the 3' end of the RTG and wherein the HSG-TBL is separated from the RDG by 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides; wherein the target site sequence is a polynucleotide sequence; and / or wherein the donor binding loop of the bridgeRNA comprises:a left-donor guide (LDG) comprising, in the 5' to 3' direction, a nucleotide sequence complementary to a first strand of a donor site sequence wherein the 3' end of the LDG complementary to at least one of the nucleotides of a core sequence on the first strand of the donor site sequence; a right-donor guide (RDG) comprising, in the 5' to 3' direction, a nucleotide sequence that is a reverse complement to an opposite strand of the first strand of the donor site sequence wherein the 3' end of the RDG is reverse complementary to at least one of the nucleotides of the core sequence on the opposite strand to the first strand of the donor site sequence; and a handshake guide (HSG-DBL), comprising, in the 5' to 3' direction, a nucleotide sequence that is partially or fully complementary to the opposite strand of the first strand of the target site sequence (anti-HSG-target site sequence) and wherein optionally the HSG-DBL is located 3' to the 3' end of the RDG or wherein the HSG-DBL is located 3' of the 3' end of the RDG and wherein the HSG-DBL is separated from the RDG by 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides; wherein the donor site sequence is a polynucleotide sequence, and the core sequence of the donor site sequence is the same as the core sequence of the target site sequence.

43. The recombinant nucleic acid editing system of claims 12-41, wherein the target binding loop of the bridgeRNA comprises: a left-target guide (LTG) comprising, in the 5' to 3' direction, a nucleotide sequence is reverse complementary to an opposite strand to a first strand of a target site sequence wherein the 3' end of the LTG is complementary to at least one of the nucleotides of a core sequence on the opposite strand to the first strand of the target site sequence; a right-target guide (RTG) comprising, in the 5' to 3' direction, a nucleotide sequence that is a complementary to the first strand of the target site sequence wherein the 3' end of the RTG is complementary to at least one of the nucleotides of the core sequence on the first strand of the target site sequence; and a handshake guide (HSG-TBL) comprising, in the 5' to 3' direction, a nucleotide sequence that is partially or fully complementary to the first strand of the donor site sequence (the anti- HSG-donor site sequence) and wherein optionally the HSG-TBL is located 3' to the 3' end ofthe RTG or wherein the HSG-TBL is located 3' of the 3' end of the RTG and wherein theHSG-TBL is separated from the RTG by 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides; wherein the target site sequence is a polynucleotide sequence; and / or wherein the donor binding loop of the bridgeRNA comprises: a left-donor guide (LDG) comprising, in the 5' to 3' direction, a nucleotide sequence complementary to a first strand of a donor site sequence wherein the 3' end of the LDG complementary to at least one of the nucleotides of a core sequence on the first strand of the donor site sequence; a right-donor guide (RDG) comprising, in the 5' to 3' direction, a nucleotide sequence that is a reverse complement to an opposite strand of the first strand of the donor site sequence wherein the 3' end of the RDG is reverse complementary to at least one of the nucleotides of the core sequence on the opposite strand to the first strand of the donor site sequence; and a handshake guide (HSG-DBL), comprising, in the 5' to 3' direction, a nucleotide sequence that is partially or fully complementary to the first strand of the target site sequence (the anti- HSG-target site sequence) and wherein optionally the HSG-DBL is located 3' to the 3' end of the RDG or wherein the HSG-DBL is located 3' of the 3' end of the RDG and wherein the HSG-DBL is separated from the RDG by 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides; wherein the donor site sequence is a polynucleotide sequence, and the core sequence of the donor site sequence is the same as the core sequence of the target site sequence.

44. The recombinant nucleic acid editing system of claim 42, wherein the HSG-TBL is partially complementary or fully complementary to one or more nucleotides immediately 5' to the core sequence on the opposite strand of the donor site sequence or wherein the HSG- TBL is partially complementary or fully complementary to one or more nucleotides located 5' to the core sequence on the opposite strand of the donor site sequence and wherein the one or more nucleotides are separated from the core sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides.

45. The recombinant nucleic acid editing system of claim 43, wherein the HSG-TBL is partially complementary or fully complementary to one or more nucleotides immediately 5' to the core sequence on the first strand of the donor site sequence or wherein the HSG-TBL is partially complementary or fully complementary to one or more nucleotides located 5' to the core sequence on the opposite strand of the donor site sequence and wherein the one or more nucleotides are separated from the core sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides.

46. The recombinant nucleic acid editing system of claim 42, wherein the HSG-DBL is partially complementary or fully complementary to one or more nucleotides immediately 5' to the core sequence on the opposite strand of the target site sequence or wherein the HSG- DBL is partially complementary or fully complementary to one or more nucleotides located 5' to the core sequence on the opposite strand of the target site sequence and wherein the one or more nucleotides are separated from the core sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides.

47. The recombinant nucleic acid editing system of claim 43, wherein the HSG-DBL is partially complementary or fully complementary to one or more nucleotides immediately 5' to the core sequence on the first strand of the target site sequence or wherein the HSG-DBL is partially complementary or fully complementary to one or more nucleotides located 5' to the core sequence on the opposite strand of the target site sequence and wherein the one or more nucleotides are separated from the core sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides.

48. The recombinant nucleic acid editing system of claims 42 or 43, wherein the target site sequence comprises sequence X1X2X3X4X5X6X7X8X9X10X11X12X13X14 where X is any nucleotide, and XsX9 are the core, one or more of X1X2X3X4X5X6X7 are the anti-HSG- target site sequence, optionally wherein Xe and X7 are the anti-HSG-target site sequence, and one or more of X12, X13, and X14 are optionally part of the target site sequence.

49. The recombinant nucleic acid editing system of claim 48, wherein the bridgeRNA encodes an LTG in the 5' to 3' direction X1X2X3X4X5X6X7X8 or X1X2X3X4X5X6X7X8X9 and an RTG in the 5' to 3' direction Y14Y13Y12Y11Y10Y9Y8 or Y14Y13Y12Y11Y10Y9 where Y is the complementary nucleotide to X and one or more of Y 14, Y 13, and Y 12 are optionally part of RTG, and a HSG-TBL in the 5' to 3' direction Zi or Z1Z2 wherein the HSG-TBL iscomplementary to a portion of the donor site sequence and does not base pair with the target site sequence.

50. The recombinant nucleic acid editing system of claims 42 or 43, wherein the donor site sequence comprises sequence STfR-Xni-XiX2X3X4X5X6X7XsX9XioXiiXi2Xi3Xi4- Xn2-STIR where X is any nucleotide, one or more of X12, X13, and X14 are optionally part of the donor site sequence, STIR is optional, but if present is a sub-terminal inverted repeat comprising 2 to 20 nucleotides, and XsX9 are the core, one or more of X1X2X3X4X5X6X7 are the anti-HSG-donor site sequence, optionally wherein XeX? is the anti-HSG-donor site sequence, and nl and n2 can independently be zero to 10.

51. The recombinant nucleic acid editing system of claim 50, wherein the bridgeRNA encodes an LDG in the 5' to 3' direction X1X2X3X4X5X6X7X8 orXiX2X3X4XsX6X7X8X9 and an RDG the 5' to 3' direction Y14Y13Y12Y11Y10Y9Y8 or Y14Y13Y12Y11Y10Y9 where Y is the complementary nucleotide to X and one or more of Y 14, Y 13, and Y 12 are optionally part of RDG, and a HSG-DBL in the 5' to 3' direction Zi or Z1Z2 wherein the HSG-DBL is complementary to a portion of the target site sequence and does not base pair with the donor site sequence.

52. The recombinant nucleic acid editing system of any of the preceding claims, wherein the bridgeRNA is engineered, wherein the bridgeRNA comprises one or more engineered HSGs, optionally wherein the bridgeRNA further comprises a wildtype (nonengineered) HSG.

53. The recombinant nucleic acid editing system of claim 52, wherein the nucleotide sequence of the HSG-DBL is engineered to alter its base pairing with said target and / or donor nucleic acid sequences, optionally wherein the engineered HSG-DBL has increased complementarity with the anti-HSG-target site sequence, and optionally wherein the engineered HSG-DBL has decreased complementarity with the anti-HSG-donor site sequence.

54. The recombinant nucleic acid editing system of any of claims 52 or 53, wherein the nucleotide sequence of the HSG-TBL is engineered to alter its base pairing with said target and / or donor nucleic acid sequences, optionally wherein the engineered HSG-TBL has increased complementarity with the anti-HSG-donor site sequence, and optionally whereinthe engineered HSG-TBL has decreased complementarity with the anti-HSG-target site sequence.

55. The recombinant nucleic acid editing system of any of claims 52-54, wherein the one or more engineered HSGs with has an altered interaction between the HSGs and the anti- HSG site sequences, wherein the base pairing between the HSG(s) and the target and / or donor sequence sites may mediate recombination, and / or wherein the one or more engineered HSG(s) possess decreased or no base pairing with between the HSG(s) and the target and / or donor sequence sites before strand exchange.

56. The recombinant nucleic acid editing system of any of claims 52-55, wherein the engineered HSG is a HSG-DBL, wherein the HSG-DBL has reduced or abolished pre-strand exchange locking between the HSG-DBL and the anti-HSG-donor site sequence.

57. The recombinant nucleic acid editing system of any of claims 52-56, wherein the engineered HSG is a HSG-TBL, wherein the HSG-TBL has a reduced or abolished pre-strand exchange locking between the HSG-TBL and the anti-HSG-target site sequence.

58. The recombinant nucleic acid editing system of any of claims 52-57, wherein the length of one or more HSG(s) is engineered, wherein the length(s) of the HSG(s) is reduced or increased.

59. The recombinant nucleic acid editing system of any of claims 58, wherein the HSG-TBL is extended, and wherein one or more nucleotides to the 3’ end of the HSG-TBL are modified to base pair with the nucleotides to the 5’ end of the anti-HSG-donor site sequence.

60. The recombinant nucleic acid editing system of any of claims 58-59, wherein the HSG-DBL is extended, and wherein one or more nucleotides to the 3’ end of the HSG-DBL are modified to base pair with the nucleotides to the 5’ end of the anti-HSG-target site sequence.

61. The recombinant nucleic acid editing system of any of claims 58-60, wherein the length(s) of the HSG(s) are increased as a result of substitution, insertion, deletion, or any combination thereof of one or more nucleotides to the 3’ end of the HSG.

62. The recombinant nucleic acid editing system of any of claims 52-61, wherein the one or more nucleotides to the 3’ end of the HSG are substituted to nucleotides capable of base pairing with the nucleotides to the 5’ end of the post-exchange anti-HSG site sequence.

63. The recombinant nucleic acid editing system of any of claims 52-62, wherein the one or more nucleotides to the 3’ end of the HSG-DBL are substituted to nucleotides capable of base pairing with the nucleotides to the 5’ end of the anti-HSG-target site sequence.

64. The recombinant nucleic acid editing system of any of claims 52-63, wherein the one or more nucleotides to the 3’ end of the HSG-TBL are substituted to nucleotides capable of base pairing with the nucleotides to the 5’ end of the anti-HSG-donor site sequence.

65. The recombinant nucleic acid editing system of any of claims 52-64, wherein one or more nucleotides are inserted adjacent to or near the 3’ end of the HSG.

66. The recombinant nucleic acid editing system of claim 65, wherein the inserted nucleotides are capable of base pairing with the nucleotides to the 5’ end of the postexchange anti-HSG site sequence, and / or wherein the inserted nucleotides reduce binding between the HSG and the nucleotides to the 5’ end of the pre-exchange anti-HSG site sequence.

67. The recombinant nucleic acid editing system of claim 65, wherein the one or more nucleotides are inserted adjacent to or near the 3’ end of the HSG-DBL, and wherein the inserted nucleotides are capable of base pairing with the nucleotides to the 5’ end of the anti- HSG-target site sequence.

68. The recombinant nucleic acid editing system of any of claims 65 or 67, wherein the one or more nucleotides are inserted adjacent to or near the 3’ end of the HSG-DBL, and wherein the inserted nucleotides are capable of reducing base pairing between the nucleotides to the 5’ end of the anti-HSG-donor site sequence.

69. The recombinant nucleic acid editing system of claim 65, wherein the one or more nucleotides are inserted adjacent to or near the 3’ end of the HSG-TBL, and wherein the inserted nucleotides are capable of base pairing with the nucleotides to the 5’ end of the anti- HSG-donor site sequence.

70. The recombinant nucleic acid editing system of any of claims 65 or 69, wherein one or more nucleotides are inserted adjacent to or near the 3’ end of the HSG-TBL, and wherein the inserted nucleotides are capable of reducing base pairing between the nucleotides to the 5’ end of the anti-HSG-target site sequence.

71. The recombinant nucleic acid editing system of any of claims 52-70, wherein one or more nucleotides are deleted adjacent to or near the 3’ end of the HSG.

72. The recombinant nucleic acid editing system of claim 71, wherein the nucleotides adjacent to the one or more deleted nucleotides are capable of base pairing with the nucleotides to the 5’ end of the post-exchange anti -HSG site sequence, and / or wherein the nucleotides adjacent to the one or more deleted nucleotides reduce binding between the HSG and the nucleotides to the 5’ end of the pre-exchange anti -HSG site sequence.

73. The recombinant nucleic acid editing system of claim 70, wherein the one or more nucleotides are deleted adjacent to or near the 3’ end of the HSG-DBL, and wherein the nucleotides adjacent to the one or more deleted nucleotides are capable of base pairing with the nucleotides to the 5’ end of the anti-HSG-target site sequence.

74. The recombinant nucleic acid editing system of any of claims 70 and 73, wherein the one or more nucleotides are deleted adjacent to or near the 3’ end of the HSG-DBL, and wherein the nucleotides adjacent to the one or more deleted nucleotides are capable of reducing base pairing between the nucleotides to the 5’ end of the anti-HSG-donor site sequence.

75. The recombinant nucleic acid editing system of claim 70, wherein, the one or more nucleotides are deleted adjacent to or near the 3’ end of the HSG-TBL, and wherein the nucleotides adjacent to the one or more deleted nucleotides are capable of base pairing with the nucleotides to the 5’ end of the anti-HSG-donor site sequence.

76. The recombinant nucleic acid editing system of any of claims 70 or 75, wherein the one or more nucleotides are deleted adjacent to or near the 3’ end of the HSG-TBL, and wherein the nucleotides adjacent to the one or more deleted nucleotides are capable of reducing base pairing between the nucleotides to the 5’ end of the anti-HSG-target site sequence.

77. The recombinant nucleic acid editing system of any of the preceding claims, wherein the site sequence(s) is modified.

78. The recombinant nucleic acid editing system of claim 77, wherein a target site sequence is engineered, and wherein the engineered target site sequence displays improved recombination efficiency with a bridgeRNA comprising one or more wildtype HSGs.

79. The recombinant nucleic acid editing system of any of claims 77-78, wherein a donor site sequence is engineered, and wherein the engineered donor site sequence displays improved recombination efficiency with a bridgeRNA comprising one or more wildtype HSGs.

80. The recombinant nucleic acid editing system of claims 12-79, wherein the target site sequence is located on genomic DNA, a linear dsDNA, a dsDNA plasmid, ssDNA or RNA.

81. The recombinant nucleic acid editing system of claims 13-79, wherein the donor site sequence is located on genomic DNA, a linear dsDNA, a dsDNA plasmid, ssDNA or RNA.

82. The recombinant nucleic acid editing system of claims 80-81, wherein the dsDNA plasmid further comprises a polynucleotide sequence for insertion into the donor site sequence.

83. The recombinant nucleic acid editing system of claims 80-81, wherein the dsDNA plasmid further comprises a polynucleotide sequence for insertion into the target site sequence.

84. The recombinant nucleic acid editing system of claims 80-81, wherein the target site sequence and donor site sequence on the genomic DNA are located on the same DNA strand.

85. The recombinant nucleic acid editing system of claims 80-81, wherein the target site sequence and donor site sequence on the genomic DNA are located on different chromosomes.

86. The recombinant nucleic acid editing system of claims 80-81, where target site sequence and donor site sequence comprise the same sequence.

87. The recombinant nucleic acid editing system of claims 80-81, where target site sequence and donor site sequence comprise different sequences.

88. The recombinant nucleic acid editing system of claims 12-87, wherein the bridgeRNA is a split bridgeRNA.

89. The recombinant nucleic acid editing system of claim 88, wherein the bridgeRNA comprises a first RNA molecule that comprises a first portion of the bridgeRNA and a second RNA molecule that comprises a second portion of the bridgeRNA.

90. The recombinant nucleic acid editing system of claim 89, wherein the first portion of the bridgeRNA comprises the internal loop comprising the first nucleotide sequence that is complementary to the first target site sequence of the target DNA, the second nucleotide sequence that is complementary to a second target site sequence and the handshake guide nucleotide sequence that is complementary to a donor site sequence and the second portion of the bridgeRNA comprises the second internal loop comprising the third nucleotide sequence that is complementary to the first donor site sequence of a donor DNA, and the fourth nucleotide sequence that is complementary to the second donor site sequence, and the handshake guide nucleotide sequence that is complementary to a target site sequence.

91. The recombinant nucleic acid editing system of claim 89-90, wherein the first portion of the bridgeRNA is encoded on a nucleic acid and is operably linked to a first promoter and the second portion of the bridgeRNA is encoded on the same or a different nucleic acid and is operably linked to a second promoter.

92. The recombinant nucleic acid editing system of claim 89-90, wherein the first portion of the bridgeRNA and second portion are encoded on a nucleic acid and further comprise one or more ribozyme sites which results in cleavage of the expressed bridgeRNA into the first and second RNA molecules.

93. The recombinant nucleic acid editing system of claims 42-92, wherein any of the LTG, RTG, LDG, and / or RDG of the bridgeRNA are complementary to more nucleotides intheir respective target site sequence or donor site sequence than the number of complementary nucleotides in its corresponding naturally occurring bridgeRNA.

94. The recombinant nucleic acid editing system of claim 93, wherein the total number of nucleotides comprising the target binding loop and / or donor binding loop of the bridgeRNA is the same as its corresponding naturally occurring bridgeRNA.

95. The recombinant nucleic acid editing system of claim 93, wherein the total number of nucleotides comprising the target binding loop and / or donor binding loop of the bridgeRNA is increased as compared to its corresponding naturally occurring bridgeRNA.

96. The recombinant nucleic acid editing system of claims 42-92, wherein any of the LTG, RTG, LDG, and / or RDG of the bridgeRNA are not complementary to a nucleotide of the core sequence in their respective target site sequence or donor site sequence but are complementary to the same number of nucleotides in the respective target site sequence or donor site sequence as its corresponding naturally occurring bridgeRNA.

97. The recombinant nucleic acid editing system of claims 42 or 43, wherein the target site sequence comprises sequence X-1X1X2X3X4X5X6X7X8X9X10X11X12X13X14 where X is any nucleotide, and XsX9 are the core and one or more of X12, X13, and X14 are optionally part of the target site sequence and wherein the bridgeRNA encodes an LTG in the 5' to 3' direction X-1X1X2X3X4X5X6X7 and an RTG in the 5' to 3' direction Y14Y13Y12Y11Y10Y9Y8 where Y is the complementary nucleotide to X and one or more of Y14, Y13, and Y12 are optionally part of RTG, and a HSG-TBL in the 5' to 3' direction Zi or Z1Z2 wherein the HSG-TBL is complementary to a portion of the donor site sequence and does not base pair with the target site sequence.

98. The recombinant nucleic acid editing system of claims 42 or 43, wherein the donor site sequence comprises sequence STIR-Xni-XiX2X3X4X5X6X7XsX9XioXnXi2Xi3Xi4- Xn2-STIR where X is any nucleotide, one or more of X12, X13, and X14 are optionally part of the donor site sequence, STIR is optional, but if present is a sub-terminal inverted repeat comprising 2 to 20 nucleotides, and XsX9 are the core, and nl and n2 can independently be 1 to 10, and wherein the bridgeRNA encodes an LDG in the 5' to 3' direction X- 1X1X2X3X4X5X6X7 and an RDG in the 5' to 3' direction Y14Y13Y12Y11Y10Y9Y8 where Y is the complementary nucleotide to X and one or more of Y14, Y13, and Y12 are optionally part ofRDG, and a HSG-DBL in the 5' to 3' direction Zi or Z1Z2 wherein the HSG-DBL is complementary to a portion of the target site sequence and does not base pair with the donor site sequence.

99. The recombinant nucleic acid editing system of claims 42-92, wherein one or more of the nucleotides of the LTG, RTG, LDG, RDG, HSG-TBL, and / or HSG-DBL of the bridgeRNA base pair to a nucleotide of their respective target site sequence or donor site sequence via non-canonical base pairing.

100. The recombinant nucleic acid editing system of any of the preceding claims, wherein the bridgeRNA comprises a scaffold, wherein the scaffold is engineered, wherein the scaffold sequence portion excludes HSGs, RTG, LTG, LDG, and RDG portions, and optionally wherein the bridgeRNA comprising an engineered scaffold is capable of increased bridgeRNA expression, activity, stability, efficiency specificity, and / or imparts novel activities.

101. The recombinant nucleic acid editing system of claim 100, wherein the engineered scaffold comprises an engineered ancillary sequence, and wherein the engineered ancillary sequence is shortened, extended, truncated, and / or modified.

102. The recombinant nucleic acid editing system of any of claims 100-101, wherein the engineered scaffold comprises an engineered TBL stem, and wherein the engineered TBL stem is shortened, extended, truncated, and / or modified.

103. The recombinant nucleic acid editing system of any of claims 100-102, wherein the engineered scaffold comprises an engineered DBL stem, and wherein the engineered DBL stem is shortened, extended, truncated, and / or modified.

104. The recombinant nucleic acid editing system of any of claims 100-103, wherein the engineered scaffold comprises an engineered stem, and wherein one or more nucleotides of the engineered stem are modified.

105. The recombinant nucleic acid editing system of any of claims 100-104, wherein the engineered stem provides enhanced stability of binding to the recombinase, enhanced tetramerization, and / or increased recombination efficiency.

106. The recombinant nucleic acid editing system of any of claims 100-105, wherein the modification of the engineered stem nucleotides comprises one or more nucleotides inserted into the non-engineered stem sequence, and wherein the one or more inserted nucleotides are capable of increasing the number of base pairs.

107. The recombinant nucleic acid editing system of any of claims 100-106, wherein the modification of the engineered stem nucleotides comprises one or more substituted nucleotides, wherein the substituted nucleotides are capable of increasing the strength of base pairing, and optionally wherein an A:T base pair in a non-engineered stem is replaced with a G:C base pair in the engineered stem.

108. The recombinant nucleic acid editing system of any of claims 100-107, wherein the engineered scaffold removes a bulge from a non-engineered scaffold, and optionally wherein the engineered scaffold comprises an engineered stem as a result of removing the bulge.

109. The recombinant nucleic acid editing system of any of claims 100-108, wherein the engineered scaffold comprises an engineered TBL, and wherein one or more of the TBL stems is engineered to be shortened, extended, truncated, and / or modified.

110. The recombinant nucleic acid editing system of any of claims 100-109, wherein the engineered scaffold comprises an engineered DBL, and wherein one or more of the DBL stems is engineered to be shortened, extended, truncated, and / or modified.

111. The recombinant nucleic acid editing system of claim 110, wherein the engineered scaffold comprises an engineered loop of the DBL stem loop, wherein one or more nucleotides of the loop of the engineered DBL stem loop are modified, and optionally wherein the engineered loop of the DBL stem loop provides enhanced stability of binding to the recombinase, enhanced tetramerization, and / or increased recombination efficiency.

112. The recombinant nucleic acid editing system of claim 109, wherein the stem loop of the TBL is replaced with the stem loop sequence of the DBL.

113. The recombinant nucleic acid editing system of claim 109, wherein the stem loop of the DBL is replaced with the stem loop sequence of the TBL.

114. The recombinant nucleic acid editing system of any of claims 100-113, wherein the engineered scaffold of the engineered bridgeRNA does not comprise a component present in the non-engineered scaffold, and optionally wherein the engineered bridgeRNA does not comprise a UU bulge present in the non-engineered bridgeRNA.

115. The recombinant nucleic acid editing system of any of claims 100-114, wherein the engineered scaffold comprises one or more of the ends of a bridgeRNA, and wherein the bridgeRNA is a split-bridgeRNA.

116. The recombinant nucleic acid editing system of any of claims 100-115, wherein one or more of the ends of the bridgeRNA are modified chemically or biochemically.

117. The recombinant nucleic acid editing system of any of claims 100-116, wherein the engineered scaffold comprises one or more nucleotides comprising a protecting group.

118. The recombinant nucleic acid editing system of claim 117, wherein the protecting group is a pseudoknot, an mRNA cap, and a tail.

119. The recombinant nucleic acid editing system of any of claims 100-118, wherein the engineered scaffold comprises one or more DNA nucleotides.

120. The recombinant nucleic acid editing system of any of claims 100-119, wherein the engineered scaffold comprises one or more non-canonical nucleotides, and optionally wherein the non-canonical nucleotide is inosine.

121. The recombinant nucleic acid editing system of any of claims 100-120, wherein the engineered scaffold comprises chemical modifications, and wherein the chemical modifications are within the bridgeRNA, at one or more ends of the bridgeRNA, and / or a base modification.

122. The recombinant nucleic acid editing system of claim 121, wherein the chemical modification is a 2’-3’ cyclic phosphate.

123. The recombinant nucleic acid editing system of any of claims 100-122, the bridgeRNA is a chimeric bridgeRNA, and optionally, wherein the chimeric bridgeRNA comprises a first component from a first IS110 ortholog and a second component from a second IS110 ortholog.

124. The recombinant nucleic acid editing system of any of claims 100-123, the bridgeRNA is circular or a portion of the bridgeRNA is circular.

125. The recombinant nucleic acid editing system of claim 124, wherein the TBL of the bridgeRNA is circular.

126. The recombinant nucleic acid editing system of claim 124, wherein the DBL of the bridgeRNA is circular.

127. The recombinant nucleic acid editing system of any of claims 124-126, the circular bridgeRNA or circular portion of the bridgeRNA is the result of circularization by a Twister sister system or a hammerhead ribozyme system.

128. The recombinant nucleic acid editing system of any of the preceding claims, wherein the engineered bridgeRNA recombines a first donor DNA molecule, a second donor DNA molecule, and a target and donor DNA molecule at a rate different from the rate of recombination by a non-engineered bridgeRNA, and wherein the engineered bridgeRNA and the non-engineered bridgeRNA comprise the same HSG(s).

129. The recombinant nucleic acid editing system of claim 128, wherein the rate of recombination for the engineered bridgeRNA is greater than the rate of the non-engineered bridgeRNA.

130. The recombinant nucleic acid editing system of claim 128, wherein the rate of recombination for the engineered bridgeRNA is lesser than the rate of the non-engineered bridgeRNA.

131. The recombinant nucleic acid editing system of any of the preceding claims, wherein the engineered bridgeRNA comprises a different number of TBLs compared to a non-engineered bridgeRNA, and wherein the engineered bridgeRNA and the non-engineered bridgeRNA comprise the same HSG(s).

132. The recombinant nucleic acid editing system of any of the preceding claims, wherein the engineered bridgeRNA comprises a different number of DBLs compared to a non-engineered bridgeRNA, and wherein the engineered bridgeRNA and the non-engineered bridgeRNA comprise the same HSG(s).

133. The recombinant nucleic acid editing system of any of the preceding claims, wherein the engineered bridgeRNA comprises one or more RNA aptamers.

134. The recombinant nucleic acid editing system of claim 133, wherein the RNA aptamer is a MS2 stem loop, wherein the MS2 stem loop is capable of recruiting an MCP domain, and optionally wherein the MCP domain is fused to a bridge recombinase.

135. The recombinant nucleic acid editing system of any of the preceding claims, wherein the engineered bridgeRNA has modified tetramerization of the recombination complex.

136. The recombinant nucleic acid editing system of any of the preceding claims, wherein the engineering bridgeRNA imparts modified properties to the assembly of the recombination complex.

137. A recombinant nucleic acid editing system comprising: a) an IS110 family transposase, or a nucleic acid comprising a sequence encoding the IS110 family transposase; and b) a bridgeRNA, or a nucleic acid comprising a sequence encoding the bridgeRNA, the bridgeRNA comprising at least one stem-loop structure and further comprising at least one internal donor binding loop comprising a first nucleotide sequence that is complementary to a first donor site sequence of a donor DNA, a second nucleotide sequence that is complementary to a second donor site sequence which is on the opposite strand of the donor DNA to the first donor site sequence, and a handshake guide nucleotide sequence that is complementary to a target site sequence, and wherein the bridgeRNA is capable of forming a complex with the IS110 family transposase.

138. A recombinant nucleic acid editing system comprising: a) an IS110 family transposase, or a nucleic acid comprising a sequence encoding the IS110 family transposase; and b) a bridgeRNA, or a nucleic acid comprising a sequence encoding the bridgeRNA, the bridgeRNA comprising at least one stem-loop structure and further comprising at least one internal donor binding loop comprising a first nucleotide sequence that is complementary to afirst donor site sequence of a donor DNA, a second nucleotide sequence that is complementary to a second donor site sequence which is on the opposite strand of the donor DNA to the first donor site sequence, and a handshake guide nucleotide sequence that is complementary to a sequence of a second donor DNA, and wherein the bridgeRNA is capable of forming a complex with the IS110 family transposase.

139. The recombinant nucleic acid editing system of claim 137-138, wherein the bridgeRNA does not comprise a target binding loop.

140. The recombinant nucleic acid editing system of claim 137-138, wherein the bridgeRNA comprises a second internal loop comprising a target binding loop, wherein the target binding loop does not bind to any nucleic acids of the nucleic acid editing system.

141. The recombinant nucleic acid editing system of claims 137-138, wherein the IS110 family transposase comprises a RuvC-like DEDD catalytic domain and a transposase domain.

142. The recombinant nucleic acid editing system of claim 141, wherein the IS110 family transposase further comprises a linker domain between the RuvC-like DEDD catalytic domain and transposase domain.

143. The recombinant nucleic acid editing system of claim 142, wherein the linker domain comprises a coiled-coil linker domain.

144. The recombinant nucleic acid editing system of claims 141-143, wherein the RuvC-like DEDD catalytic domain comprises an amino acid sequence at least 50% identical to a RuvC-like DEDD catalytic domain sequence of SEQ ID NOs: 10176-10523 or 10524- 20350.

145. The recombinant nucleic acid editing system of claims 141-143, wherein theIS110 family transposase comprises a RuvC-like DEDD catalytic domain that forms a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621.

146. The recombinant nucleic acid editing system of claim 145, wherein the IS110 family transposase RuvC-like DEDD catalytic domain comprises a tertiary structure similar to a tertiary structure of the RuvC-like DEDD catalytic domain of IS621 if the templatemodeling score (TM-score) for the RuvC-like DEDD catalytic domain of the IS110 family transposase is 0.5 or higher.

147. The recombinant nucleic acid editing system of claim 146, wherein the RuvC-like DEDD catalytic domain comprises an amino acid sequence at least 15% identical to a RuvC- like DEDD catalytic domain sequence of SEQ ID NOs: 10176-10523 or 10524-20350.

148. The recombinant nucleic acid editing system of claims 141-143, wherein the transposase domain comprises an amino acid sequence at least 50% identical to a transposase domain sequence of SEQ ID NOs: 10176-10523 or 10524-20350.

149. The recombinant nucleic acid editing system of claims 141-143, wherein the IS110 family transposase comprises a transposase domain that forms a similar tertiary structure to the transposase domain of IS621.

150. The recombinant nucleic acid editing system of claim 149, wherein the IS110 family transposase domain comprises a tertiary structure similar to a tertiary structure of the transposase domain of IS621 if the template modeling score (TM-score) for the transposase domain of the IS110 family transposase is 0.5 or higher.

151. The recombinant nucleic acid editing system of claim 150, wherein the transposase domain comprises an amino acid sequence at least 15% identical to a transposase domain sequence of SEQ ID NOs: 10176-10523 or 10524-20350.

152. The recombinant nucleic acid editing system of claims 141-143, wherein theIS110 family transposase comprises an amino acid sequence at least 50% identical to SEQ ID NOs: 10176-10523 or 10524-20350.

153. The recombinant nucleic acid editing system of claims 141-143, wherein the IS110 family transposase further comprises a tertiary structure similar to a tertiary structure of IS621.

154. The recombinant nucleic acid editing system of claim 153, wherein the IS110 family transposase comprises a tertiary structure similar to a tertiary structure of IS621 if the template modeling score (TM-score) for the transposase is 0.5 or higher.

155. The recombinant nucleic acid editing system of claim 154, wherein the transposase domain comprises an amino acid sequence at least 15% identical to a transposase domain sequence of SEQ ID NOs: 10176-10523 or 10524-20350.

156. The recombinant nucleic acid editing system of claims 137-155, wherein the IS110 family transposase is an IS 110 group transposase.

157. The recombinant nucleic acid editing system of claims 137-155, wherein the IS110 family transposase is an IS1111 group transposase.

158. The recombinant nucleic acid editing system of claim 157, wherein the IS1111 group transposase is IS1111 A or IS1111_229727.

159. The recombinant nucleic acid editing system of claim 156, wherein the IS110 group transposase is IS621, ISPal l, IsPa29, ISMmgl, ISPfll, ISMae40, ISStma6, ISAzs32, ISMex9, ISCARN28, IS Aar 16, ISCps7, ISPpu9, ISRel9, ISEsa2, ISMma5, IS900, or ISHne5.

160. The recombinant nucleic acid editing system of claim 156, wherein the IS110 group transposase comprises an amino acid sequence at least 50% identical to IS621 (SEQ ID NO: 10176) or IS621_23122 (SEQ ID NO: 10181).

161. The recombinant nucleic acid editing system of claims 137-160, wherein the stem of the stem-loop is 5 to 35 nucleotides long and the loop is 3-10 nucleotides long, and the donor binding loop is 5 to 20 nucleotides long.

162. The recombinant nucleic acid editing system of claim 161, wherein the stem of the second stem-loop structure comprises 1 to 4 loops or bubbles that are each 1 to 10 nucleotides long.

163. The recombinant nucleic acid editing system of claims 137-162, wherein the donor binding loop of the bridgeRNA comprises: a left-donor guide (LDG) comprising, in the 5' to 3' direction, a nucleotide sequence complementary to a first strand of a donor site sequence wherein the 3' end of the LDG complementary to at least one of the nucleotides of a core sequence on the first strand of the donor site sequence;a right-donor guide (RDG) comprising, in the 5' to 3' direction, a nucleotide sequence that is a reverse complement to an opposite strand of the first strand of the donor site sequence wherein the 3' end of the RDG is reverse complementary to at least one of the nucleotides of the core sequence on the opposite strand to the first strand of the donor site sequence; and a handshake guide (HSG-DBL), comprising, in the 5' to 3' direction, a nucleotide sequence that is partially or fully complementary to the opposite strand of the first strand of a target site sequence (anti-HSG-target site sequence) and wherein optionally the HSG-DBL is located 3' to the 3' end of the RDG or wherein the HSG-DBL is located 3' of the 3' end of the RDG and wherein the HSG-DBL is separated from the RDG by 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides; wherein the donor site sequence is a polynucleotide sequence, and the core sequence of the donor site sequence is the same as the core sequence of a target site sequence.

164. The recombinant nucleic acid editing system of claims 137-162, wherein the donor binding loop of the bridgeRNA comprises: a left-donor guide (LDG) comprising, in the 5' to 3' direction, a nucleotide sequence complementary to a first strand of a donor site sequence wherein the 3' end of the LDG complementary to at least one of the nucleotides of a core sequence on the first strand of the donor site sequence; a right-donor guide (RDG) comprising, in the 5' to 3' direction, a nucleotide sequence that is a reverse complement to an opposite strand of the first strand of the donor site sequence wherein the 3' end of the RDG is reverse complementary to at least one of the nucleotides of the core sequence on the opposite strand to the first strand of the donor site sequence; and a handshake guide (HSG-DBL), comprising, in the 5' to 3' direction, a nucleotide sequence that is partially or fully complementary to the first strand of a target site sequence (anti-HSG- target site sequence) and wherein optionally the HSG-DBL is located 3' to the 3' end of the RDG or wherein the HSG-DBL is located 3' of the 3' end of the RDG and wherein the HSG- DBL is separated from the RDG by 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides; wherein the donor site sequence is a polynucleotide sequence, and the core sequence of the donor site sequence is the same as the core sequence of a target site sequence.

165. The recombinant nucleic acid editing system of claims 137-162, wherein the donor binding loop of the bridgeRNA comprises: a left-donor guide (LDG) comprising, in the 5' to 3' direction, a nucleotide sequence complementary to a first strand of a donor site sequence of a first donor DNA wherein the 3' end of the LDG complementary to at least one of the nucleotides of a core sequence on the first strand of the donor site sequence of the first donor DNA; a right-donor guide (RDG) comprising, in the 5' to 3' direction, a nucleotide sequence that is a reverse complement to an opposite strand of the first strand of the first donor site sequence of the first donor DNA wherein the 3' end of the RDG is reverse complementary to at least one of the nucleotides of the core sequence on the opposite strand to the first strand of the donor site sequence of the first donor DNA; and a handshake guide (HSG-DBL), comprising, in the 5' to 3' direction, a nucleotide sequence that is partially or fully complementary to the opposite strand of the first strand of a donor site sequence of a second donor DNA and wherein optionally the HSG-DBL is located 3' to the 3' end of the RDG or wherein the HSG-DBL is located 3' of the 3' end of the RDG and wherein the HSG-DBL is separated from the RDG by 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides; wherein the donor site sequence of the first donor DNA is a polynucleotide sequence, the donor site sequence of the second donor DNA is a polynucleotide sequence, and the core sequence of the donor site sequence of the first donor DNA is the same as the core sequence of a donor site sequence of the second donor DNA.

166. The recombinant nucleic acid editing system of claims 137-162, wherein the donor binding loop of the bridgeRNA comprises: a left-donor guide (LDG) comprising, in the 5' to 3' direction, a nucleotide sequence complementary to a first strand of a donor site sequence of a first donor DNA wherein the 3' end of the LDG complementary to at least one of the nucleotides of a core sequence on the first strand of the donor site sequence of the first donor DNA; a right-donor guide (RDG) comprising, in the 5' to 3' direction, a nucleotide sequence that is a reverse complement to an opposite strand of the first strand of the first donor site sequenceof the first donor DNA wherein the 3' end of the RDG is reverse complementary to at least one of the nucleotides of the core sequence on the opposite strand to the first strand of the donor site sequence of the first donor DNA; and a handshake guide (HSG-DBL), comprising, in the 5' to 3' direction, a nucleotide sequence that is partially or fully complementary to the first strand of a donor site sequence of a second donor DNA and wherein optionally the HSG-DBL is located 3' to the 3' end of the RDG or wherein the HSG-DBL is located 3' of the 3' end of the RDG and wherein the HSG-DBL is separated from the RDG by 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides; wherein the donor site sequence of the first donor DNA is a polynucleotide sequence, the donor site sequence of the second donor DNA is a polynucleotide sequence, and the core sequence of the donor site sequence of the first donor DNA is the same as the core sequence of a donor site sequence of the second donor DNA.

167. The recombinant nucleic acid editing system of claims 163, wherein the HSG- DBL is partially complementary or fully complementary to one or more nucleotides immediately 5' to the core sequence on the opposite strand of the target site sequence or wherein the HSG-DBL is partially complementary or fully complementary to one or more nucleotides located 5' to the core sequence on the opposite strand of the target site sequence and wherein the one or more nucleotides are separated from the core sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides.

168. The recombinant nucleic acid editing system of claim 164, wherein the HSG-DBL is partially complementary or fully complementary to one or more nucleotides immediately 5' to the core sequence on the first strand of the target site sequence or wherein the HSG-DBL is partially complementary or fully complementary to one or more nucleotides located 5' to the core sequence on the opposite strand of the target site sequence and wherein the one or more nucleotides are separated from the core sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides.

169. The recombinant nucleic acid editing system of claims 165, wherein the HSG- DBL is partially complementary or fully complementary to one or more nucleotides immediately 5' to the core sequence on the opposite strand of the donor site sequence of the second donor DNA or wherein the HSG-DBL is partially complementary or fullycomplementary to one or more nucleotides located 5' to the core sequence on the opposite strand of the donor site sequence of the second donor DNA and wherein the one or more nucleotides are separated from the core sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides.

170. The recombinant nucleic acid editing system of claim 166, wherein the HSG-DBL is partially complementary or fully complementary to one or more nucleotides immediately 5' to the core sequence on the first strand of the donor site sequence of the second donor DNA or wherein the HSG-DBL is partially complementary or fully complementary to one or more nucleotides located 5' to the core sequence on the opposite strand of the donor site sequence of the second donor DNA and wherein the one or more nucleotides are separated from the core sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 intervening nucleotides.

171. The recombinant nucleic acid editing system of claims 135-170, wherein the bridgeRNA is engineered, wherein the bridgeRNA comprises one or more engineered HSGs.

172. The recombinant nucleic acid editing system of claim 171, wherein the nucleotide sequence of the HSG-DBL is engineered to alter its base pairing with said target and / or donor nucleic acid sequences.

173. The recombinant nucleic acid editing system of claims 137-172, wherein the target and / or donor site sequence(s) are located on genomic DNA, a linear dsDNA, a dsDNA plasmid, ssDNA or RNA.

174. The recombinant nucleic acid editing system of claim 173, wherein the dsDNA plasmid further comprises a polynucleotide sequence for insertion into the donor site sequence of the second donor DNA.

175. The recombinant nucleic acid editing system of claim 173, wherein the dsDNA plasmid further comprises a polynucleotide sequence for insertion into the target site sequence.

176. The recombinant nucleic acid editing system of claims 137-172, wherein the target site sequence and donor site sequence on the genomic DNA or the donor site sequence of the first donor DNA and the donor site sequence of the second donor DNA are located on the same DNA strand.

177. The recombinant nucleic acid editing system of claims 137-172, wherein the target site sequence and donor site sequence on the genomic DNA or the donor site sequence of the first donor DNA and the donor site sequence of the second donor DNA are located on different chromosomes.

178. The recombinant nucleic acid editing system of claims 137-172, where target site sequence and donor site sequence comprise the same sequence.

179. The recombinant nucleic acid editing system of claims 137-172, where target site sequence and donor site sequence comprise different sequences.

180. The recombinant nucleic acid editing system of claims 137-172, where the donor site sequence of the first donor DNA and the donor site sequence of the second donor DNA comprise the same sequence.

181. The recombinant nucleic acid editing system of claims 137-172, where the donor site sequence of the first donor DNA and the donor site sequence of the second donor DNA comprise different sequences.

182. The recombinant nucleic acid editing system of claims 137-181, wherein the bridgeRNA is a split bridgeRNA.

183. The recombinant nucleic acid editing system of claim 182, wherein the bridgeRNA comprises a first RNA molecule that comprises the internal donor binding loop comprising a first nucleotide sequence that is complementary to a first donor site sequence of a first donor DNA, a second nucleotide sequence that is complementary to a second donor site sequence which is on the opposite strand of the first donor DNA to the first donor site sequence, and a handshake guide nucleotide sequence that is complementary to a sequence of a second donor DNA; and a second RNA molecule that comprises an internal donor binding loop comprising a first nucleotide sequence that is complementary to a first donor site sequence of the second donor DNA, a second nucleotide sequence that is complementary to a second donor site sequence which is on the opposite strand of the second donor DNA to the first donor site sequence, and a handshake guide nucleotide sequence that is complementary to a sequence of the first donor DNA.

184. The recombinant nucleic acid editing system of claim 182-183, wherein the first RNA molecule of the bridgeRNA is encoded on a nucleic acid and is operably linked to a first promoter and the second RNA molecule of the bridgeRNA is encoded on the same or a different nucleic acid and is operably linked to a second promoter.

185. The recombinant nucleic acid editing system of claim 182-183, wherein the first RNA molecule of the bridgeRNA and second RNA molecule are encoded on a nucleic acid and further comprise one or more ribozyme sites which results in cleavage of the expressed bridgeRNA into the first and second RNA molecules.

186. The recombinant nucleic acid editing system of claims 165-185, wherein one or more of the nucleotides of the LDG, RDG, and / or HSG-DBL of the bridgeRNA base pair to a nucleotide of their respective donor site sequences via non-canonical base pairing.

187. The recombinant nucleic acid editing system of claims 137-186, wherein the bridgeRNA comprises a scaffold, wherein the scaffold is engineered, wherein the scaffold sequence portion excludes HSGs, RTG, LTG, LDG, and RDG portions, and optionally wherein the bridgeRNA comprising an engineered scaffold is capable of increased bridgeRNA expression, activity, stability, efficiency specificity, and / or imparts novel activities.

188. The recombinant nucleic acid editing system of any of claims 137-187, the bridgeRNA is circular or a portion of the bridgeRNA is circular.

189. The recombinant nucleic acid editing system of claims 137-187, wherein the engineered bridgeRNA recombines a first donor DNA molecule and a target DNA molecule and a first donor DNA molecule and a second donor DNA molecule.

190. The recombinant nucleic acid editing system of claims 137-187, wherein the engineered bridgeRNA recombines a first donor DNA molecule and second donor DNA molecule.

191. The recombinant nucleic acid editing system of claim 190, wherein the engineered bridgeRNA does not recombine or has reduced recombination of a target DNA molecule and a first donor DNA molecule.

192. A vector comprising any of the nucleic acids of the nucleic acid editing system of claims 1-191.

193. A host cell comprising any of the vector(s) of claim 192.

194. The recombinant nucleic acid editing system of claims 1-191, wherein any of the nucleic acids of the nucleic acid editing system further comprise an inducible promoter.

195. A method of integrating a DNA molecule of interest into a sequence specific site of a DNA of interest of a cell, the method comprising introducing into the cell: a nucleic acid editing system according to claims 1-191.

196. The method of claim 195, wherein the cell is a mammalian cell.

197. The method of claim 195, where the cell is a human cell.

198. The method of claims 195-197, wherein the DNA of interest of the cell comprises a donor site sequence and the DNA molecule of interest is part of a nucleic acid of the nucleic acid editing system comprising a target site sequence.

199. The method of claims 195-197, wherein the DNA of interest of the cell comprises a target site sequence and the DNA molecule of interest is part of a nucleic acid of the nucleic acid editing system comprising a donor site sequence.

200. The method of claim 198, wherein the DNA of interest of the cell further comprises a second donor site sequence and the DNA molecule of interest further comprises a second target site sequence and the nucleic acid editing system comprises a second bridgeRNA that targets the second donor site sequence and second target site sequence.

201. The method of claim 199, wherein the DNA of interest of the cell comprises a second target site sequence and the DNA molecule of interest further comprises a second donor site sequence and the nucleic acid editing system comprises a second bridgeRNA that targets the second donor site sequence and second target site sequence.

202. The method of claims 195-197, wherein the DNA of interest of the cell comprises a donor site sequence of a second donor DNA and the DNA molecule of interest is part of anucleic acid of the nucleic acid editing system comprising a donor site sequence of a first donor DNA.

203. The method of claims 195-197, wherein the DNA of interest of the cell comprises a donor site sequence of a donor DNA and the DNA molecule of interest is part of a nucleic acid of the nucleic acid editing system comprising a donor site sequence of a donor DNA.

204. The method of claims 198-201, wherein the sequence of the bridgeRNA was engineered before introduction of the nucleic acid editing system to bind to the donor site sequence and target site sequence.

205. The method of claims 196-204, wherein the DNA of interest of the cell is the genome of the cell.

206. The method of claims 196-204, wherein the DNA of interest of the cell is a plasmid.

207. The method of claims 196-204, wherein the donor site sequence or target site sequence of the DNA of interest of the cell has maximal orthogonality to the genome sequence of the cell.

208. The method of claim 207, wherein orthogonality is determined using a combination of Hamming distance and Levenshtein distance.

209. The method of claims 196-207, wherein the target site sequence of the DNA of interest of the cell (if present) comprises an array of target site sequences comprising repeated identical sequences.

210. The method of claims 196-207, wherein the donor site sequence of the DNA of interest of the cell (if present) comprises an array of donor site sequences comprising repeated identical sequences.

211. The method of claims 196-209, wherein the IS110 family transposase of the nucleic acid editing system is introduced to the cell as a mRNA comprising a sequence encoding the IS110 family transposase; and the bridgeRNA of the nucleic acid editing system is introduced to the cell as a RNA.

212. A method of inverting a DNA sequence of a DNA of interest of a cell, the method comprising introducing into the cell: a nucleic acid editing system according to claims 1-136, wherein a target site sequence and a donor site sequence are present on the same DNA molecule of interest and the LD of the donor site sequence and the RT of the target site sequence are on the same DNA strand, and the RD of the donor site sequence and the LT of the target site sequence are on the same strand.

213. The method of claim 212, wherein the DNA of interest of the cell is the genome of the cell.

214. The method of claims 212-213, wherein the donor site sequence is located closer to the 5’ end of the DNA molecule of interest and the target site sequence is located closer to the 3’ end of the DNA molecule of interest.

215. The method of claims 212-214, wherein the sequence of the bridgeRNA was engineered before introduction of the nucleic acid editing system to bind to the donor site sequence and target site sequence.

216. The method of claims 212-215, wherein the target site sequence of the DNA of interest of the cell comprises an array of target site sequences comprising repeated identical sequences.

217. The method of claims 212-215, wherein the donor site sequence of the DNA of interest of the cell comprises an array of donor site sequences comprising repeated identical sequences.

218. A method of inverting a DNA sequence of a DNA of interest of a cell, the method comprising introducing into the cell: a nucleic acid editing system according to claims 137- 191, wherein a first donor site sequence and a second donor site sequence are present on the same DNA molecule of interest and the LD of the first donor site sequence and the RD of the second donor site sequence are on the same DNA strand, and the RD of the first donor site sequence and the LD of the second donor site sequence are on the same strand.

219. The method of claim 218, wherein the DNA of interest of the cell is the genome of the cell.

220. The method of claims 218-219, wherein the first donor site sequence of the DNA of interest of the cell comprises an array of first donor site sequences comprising repeated identical sequences.

221. The method of claims 218-219, wherein the second donor site sequence of the DNA of interest of the cell comprises an array of second donor site sequences comprising repeated identical sequences.

222. The method of claims 218-221, wherein the sequence of the bridgeRNA was engineered before introduction of the nucleic acid editing system to bind to the first donor site sequence and second donor site sequence.

223. The method of claims 212-222, wherein the IS110 family transposase of the nucleic acid editing system is introduced to the cell as a mRNA comprising a sequence encoding the IS110 family transposase; and the bridgeRNA of the nucleic acid editing system is introduced to the cell as a RNA.

224. A method of excising a DNA sequence of a DNA of interest of a cell, the method comprising introducing into the cell: a nucleic acid editing system according to claims 1-134, wherein a target site sequence and a donor site sequence are present on the same DNA molecule of interest and the LD of the donor site sequence and the LT of the target site sequence are on the same DNA strand.

225. The method of claim 224, wherein the DNA of interest of the cell is the genome of the cell.

226. The method of claims 224-225, wherein the sequence of the bridgeRNA was engineered before introduction of the nucleic acid editing system to bind to the donor site sequence and the target site sequence.

227. The method of claims 224-226, wherein the target site sequence of the DNA of interest of the cell comprises an array of target site sequences comprising repeated identical sequences.

228. The method of claims 224-227, wherein the donor site sequence of the DNA of interest of the cell comprises an array of donor site sequences comprising repeated identical sequences.

229. A method of excising a DNA sequence of a DNA of interest of a cell, the method comprising introducing into the cell: a nucleic acid editing system according to claims 137- 191, wherein a first donor site sequence and a second donor site sequence are present on the same DNA molecule of interest and the LD of the second donor site sequence and the LD of the first donor site sequence are on the same DNA strand.

230. The method of claim 229, wherein the DNA of interest of the cell is the genome of the cell.

231. The method of claims 229-230, wherein the sequence of the bridgeRNA was engineered before introduction of the nucleic acid editing system to bind to the first donor site sequence and the second donor site sequence.

232. The method of claims 229-231, wherein the first donor site sequence of the DNA of interest of the cell comprises an array of first donor site sequences comprising repeated identical sequences.

233. The method of claims 229-232, wherein the second donor site sequence of the DNA of interest of the cell comprises an array of second donor site sequences comprising repeated identical sequences.

234. The method of claims 224-233, wherein the IS110 family transposase of the nucleic acid editing system is introduced to the cell as a mRNA comprising a sequence encoding the IS110 family transposase; and the bridgeRNA of the nucleic acid editing system is introduced to the cell as a RNA.

235. A method of translocating DNA sequences between two linear DNA molecules of interest, the method comprising introducing into a cell: a nucleic acid editing system according to claims 1-134, wherein a donor site sequence is present on a first linear DNA molecule and a target site sequence is present on a second linear DNA molecule.

236. The method of claim 235, wherein the linear DNA molecules of interest of the cell are chromosomes of the cell.

237. The method of claims 235-236, wherein the sequence of the bridgeRNA was engineered before introduction of the nucleic acid editing system to bind to the donor site sequence and the target site sequence.

238. A method of excising a repeat region of a genome of interest comprising: introducing a nucleic acid editing system of claims 1-191 wherein repeat sequences are excised from the genome of interest.

239. The method of claim 238, wherein the nucleic acid editing system comprises a bridgeRNA comprising a TBL and a DBL that target the same sequence to mediate recombination reactions between the repeats of the genome of interest.

240. The method of claim 238, wherein nucleic acid editing system comprises a bridgeRNA comprising only a DBL.

241. The method of claims 238-240, wherein the repeat region is of the FXN, HTT, CNBP, or DMPK gene.

242. The method of claims 238-241, wherein the genome of interest is the human genome.

243. The method of claims 238-242, wherein the IS110 family transposase of the nucleic acid editing system is introduced as a mRNA comprising a sequence encoding the IS110 family transposase; and the bridgeRNA of the nucleic acid editing system is introduced as a RNA.

244. A method of treating a repeat expansion disorder in a subject in need thereof comprising: administering a nucleic acid editing system of claims 1-191 wherein repeat sequences are excised from the genome of the subject.

245. The method of claim 244, wherein the nucleic acid editing system comprises abridgeRNA comprising a TBL and a DBL that target the same sequence to mediate recombination reactions between the repeats of the genome of the subject.

246. The method of claim 244, wherein nucleic acid editing system comprises a bridgeRNA comprising only a DBL.

247. The method of claims 244-246, wherein the repeat expansion disorder is Friedreich’s ataxia, Huntington’s disease, or myotonic dystrophy type 1 or type 2.

248. The method of claims 238-247, wherein the subject is a human.

249. The method of claims 244-248, wherein the IS110 family transposase of the nucleic acid editing system is administered as a mRNA comprising a sequence encoding the IS110 family transposase; and the bridgeRNA of the nucleic acid editing system is administered as a RNA.

250. A method of reverting a pre-existing inversion or translocation of a genome of interest comprising: introducing a nucleic acid editing system of claims 1-191 wherein an inversion or translocation is performed to revert a pre-existing inversion or translocation.

251. The method of claim 250, wherein the nucleic acid editing system comprises a bridgeRNA comprising a TBL and a DBL that target the breakpoints of the pre-existing inversion or translocation to mediate recombination reactions between the breakpoints of the pre-existing inversion or translocation of the genome of interest.

252. The method of claim 250, wherein the nucleic acid editing system comprises a bridgeRNA comprising a first DBL and a second DBL that target the breakpoints of the preexisting inversion or translocation to mediate recombination reactions between the breakpoints of the pre-existing inversion or translocation of the genome of interest.

253. The method of claim 250, wherein the nucleic acid editing system comprises a bridgeRNA comprising a single DBL that targets the breakpoints of the pre-existing inversionor translocation to mediate recombination reactions between the breakpoints of the inversion or translocation of the pre-existing genome of interest.

254. The method of claims 250-253, wherein the genome of interest is the human genome.

255. The method of claims 250-253, wherein the breakpoints have the same sequence.

256. The method of claims 250-253, wherein the breakpoints have different sequences.

257. The method of claims 250-256, wherein the IS110 family transposase of the nucleic acid editing system is introduced as a mRNA comprising a sequence encoding the IS110 family transposase; and the bridgeRNA of the nucleic acid editing system is introduced as a RNA.

258. A method of treating a pre-existing genomic inversion or translocation in a subject in need thereof comprising: administering a nucleic acid editing system of claims 1-191 wherein an inversion or translocation is performed to revert the pre-existing inversion or translocation of the genome of the subject.

259. The method of claim 258, wherein the nucleic acid editing system comprises a bridgeRNA comprising a TBL and a DBL that target the breakpoints of the pre-existing inversion or translocation to mediate recombination reactions between the breakpoints of the pre-existing inversion or translocation of the genome of the subject.

260. The method of claim 258, wherein the nucleic acid editing system comprises a bridgeRNA comprising a first DBL and a second DBL that target the breakpoints of the preexisting inversion or translocation to mediate recombination reactions between the breakpoints of the pre-existing inversion or translocation of the genome of the subject.

261. The method of claim 258, wherein the nucleic acid editing system comprises a bridgeRNA comprising a single DBL that targets the breakpoints of the pre-existing inversion or translocation to mediate recombination reactions between the breakpoints of the pre-existing inversion or translocation of the genome of the subject.

262. The method of claims 258-261, wherein the subject is a human.

263. The method of claims 258-261, wherein the breakpoints have the same sequence.

264. The method of claims 258-261, wherein the breakpoints have different sequences.

265. The method of claims 258-261, wherein the IS110 family transposase of the nucleic acid editing system is administered as a mRNA comprising a sequence encoding the IS110 family transposase; and the bridgeRNA of the nucleic acid editing system is administered as a RNA.

266. The nucleic acid editing system of claims 1-191, wherein the IS110 family transposase of the nucleic acid editing system in in the form of a mRNA comprising a sequence encoding the IS110 family transposase; and the bridgeRNA of the nucleic acid editing system is in the form of a RNA.

Citation Information

Patent Citations

  • Plasmids for manipulation of wolbachia

    US20220010317A1

  • Novel type vi crispr enzymes and systems

    US20230025039A1

  • Compositions comprising riboregulators and methods of use thereof

    WO2017087530A1

  • Reprogrammable ISCB nucleases and uses thereof

    WO2023097228A1

  • Programmable DNA transposases for nucleic acid manipulation

    WO2024119154A1