Programmable DNA transposases for nucleic acid manipulation

The IS110 family transposase and bridge RNA system addresses the complexity of existing DNA manipulation methods by providing a simple, reprogrammable, and efficient means for DNA sequence manipulation in cells, enabling precise DNA insertion, inversion, and excision.

JP2026500148APending Publication Date: 2026-01-06ARC RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025532045
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-09-07
Filing Date
2023-12-01
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

Existing DNA sequence manipulation methods, such as Cre-loxP and RNA-guided CRISPR effectors, require extensive protein engineering and complex delivery systems for programmable DNA insertion, inversion, and excision, lacking simplicity and ease of reprogramming.

Method used

A recombinant nucleic acid editing system utilizing an IS110 family transposase and a bridge RNA with stem-loop structures for specific DNA targeting, enabling reprogrammable DNA sequence manipulation without extensive protein engineering and easy cellular delivery.

Benefits of technology

Facilitates efficient and programmable DNA insertion, inversion, and excision with high specificity and ease of reprogramming, using a nucleic acid editing system that integrates, inverts, or excises DNA sequences in cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026500148000001_ABST
    Figure 2026500148000001_ABST
Patent Text Reader

Abstract

The present invention relates to a novel system for nucleic acid manipulation using components of the IS110 family of transposons. The IS110 transposase encoded by the IS110 component has been identified to utilize an RNA sequence called a bridge RNA. The bridge RNA targets donor and target sequence sites for polynucleotide recombination reactions. In one aspect, the present application relates to reprogramming the IS110 transposase to integrate sequences at predetermined sites using the bridge RNA. Using the IS110 transposase and bridge RNA, programmable insertion, excision recombination, and / or inversion allows the integration or transfer of any polynucleotide encoding a donor site or target site recognized by the IS110 transposase into any other polynucleotide sequence containing the target site sequence or donor site sequence, respectively. The present invention has applications in cell engineering, genome engineering, genetic medicine, synthetic biology, molecular diagnostics, transgenic organisms, and biological research.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This international patent application, entitled "INSERTION OF This application claims the benefit of and priority to U.S. application Ser. No. 63 / 385,736, entitled "PROGRAMMABLE CARGO WITH PROGRAMMABLE TRANSPOSASES," filed Sep. 7, 2023, and U.S. application Ser. No. 63 / 581,208, entitled "PROGRAMMABLE DNA TRANSPOSASES FOR NUCLEIC ACID MANIPULATION," filed Sep. 7, 2023, the contents of which are incorporated herein by reference in their entireties.

[0002] All patents, patent applications, and publications cited herein are incorporated by reference in their entireties. The disclosures of these publications are incorporated by reference in their entireties into this application.

[0003] This patent disclosure contains material that is subject to copyright protection. As the patent document or patent disclosure appears in the U.S. Patent and Trademark Office patent file or records, the copyright owner has no objection to the facsimile reproduction thereof by anyone, but otherwise reserves all copyright rights whatsoever.

[0004] Sequence Listing This application contains a "long" Sequence Listing submitted via DVD-R in lieu of a printed paper copy, which is incorporated herein by reference in its entirety. The DVD-Rs, recorded on December 1, 2023, are labeled "CRF," "Copy 1," and "Copy 2," respectively, and each contains a single, identical 819,898,112-byte (781 MB) file (2220476-00121WO1_SL.xml). [Background technology]

[0005] Background of the Invention DNA sequence manipulation—DNA insertion, inversion, and excision—is a fundamental capability underlying synthetic biology, genome engineering, cell engineering, and genetic medicine. Such reactions have been utilized for decades, such as the Cre-loxP system, which relies on the specificity of the Cre protein for predesigned loxP DNA sequences. The orientation and location of these DNA sequences relative to each other are used to achieve the insertion, inversion, or excision of long DNA sequences. More modern methods, particularly those for DNA insertion, often utilize random, semi-random, or non-programmable site-specific enzymes, such as recombinases and transposases, to manipulate DNA sequences and rely on the enzyme's encoded target sequence specificity. RNA-guided programmable nucleases, such as Cas9, have been used to excise DNA by cleaving at multiple sites, but rely on DNA repair of the resulting double-strand break. With regard to insertion and inversion, the amino acid-encoded DNA specificity of these systems is difficult to reprogram without extensive protein engineering efforts, particularly through the attachment of additional programmable targeting domains, such as RNA-guided CRISPR effectors. Other methods utilize synthetic or naturally occurring reprogrammable transposition systems that require the assembly of multiple protein and nucleic acid subunits to achieve DNA integration, but delivering these systems into cells can be complex and requires extensive engineering techniques to achieve efficient inversion or excision of DNA sequences. Therefore, there is a need for simple DNA sequence recombination systems that are reprogrammable without protein engineering and easily delivered to cells. There is a great need and potential for this. Summary of the Invention

[0006] Unless the context indicates otherwise, it is understood that any of the embodiments described below can be combined in any desired manner, and that any embodiment or combination of embodiments can be applied to each of the aspects described below.

[0007] In certain aspects, the present invention provides a recombinant nucleic acid editing system comprising: a) an IS110 family transposase or a nucleic acid comprising a sequence encoding said IS110 family transposase; and b) a nucleic acid comprising a sequence encoding a bridge RNA.

[0008] In some embodiments, the nucleic acid comprising a sequence encoding the bridge RNA comprises the left end (LE) sequence of a transposon encoding the IS110 family transposase, hi some embodiments, the nucleic acid comprising a sequence encoding the bridge RNA comprises the right end (RE) sequence of a transposon encoding the IS110 family transposase.

[0009] In some embodiments, the recombinant nucleic acid editing system further comprises a nucleic acid comprising an RE sequence and an LE sequence, or an RE sequence, a core sequence, and an LE sequence, of an IS110 element encoding the IS110 family transposase. In some embodiments, the recombinant nucleic acid editing system further comprises a nucleic acid comprising a right flanking (RF) sequence and a left flanking (LF) sequence, or an RF sequence, a core sequence, and an LF sequence, of a target site sequence of the IS110 family transposase. In some embodiments, the nucleic acid comprising the RE sequence and an LE sequence, or the RE sequence, a core sequence, and an LE sequence, further comprises a nucleic acid sequence for insertion into a target site sequence. In some embodiments, the target site sequence comprises an RF sequence and an LF sequence, or an RF sequence, a core sequence, and an LF sequence, of the IS110 family transposase. In some embodiments, the nucleic acid comprising the RF sequence and an LF sequence, or the RF sequence, a core sequence, and an LF sequence, further comprises a nucleic acid sequence for insertion into a donor site sequence. In some embodiments, the donor site sequence comprises the RE and LE sequences, or the RE, core, and LE sequences, of an IS110 element encoding the IS110 family transposase.

[0010] In some embodiments, the bridge RNA comprises a nucleotide sequence at least 50% identical to a bridge RNA sequence of SEQ ID NO: 1-348 or SEQ ID NO: 349-10175. In some embodiments, the RE sequence comprises an RE sequence of SEQ ID NO: 1-348, 30354-30529, 349-10175, or 30530-40356, the LE sequence comprises an LE sequence of SEQ ID NO: 1-348, 30354-30529, 349-10175, or 30530-40356, and / or the core sequence comprises a core sequence of SEQ ID NO: 1-348, 30354-30529, 349-10175, or 30530-40356.

[0011] In a specific aspect, the present invention provides a recombinant nucleic acid editing system comprising: a) an IS110 family transposase or a nucleic acid comprising a sequence encoding the IS110 family transposase; and b) a bridge RNA or a nucleic acid comprising a sequence encoding the bridge RNA, wherein the bridge RNA comprises at least one stem-loop structure and further comprises at least one internal loop comprising a first nucleotide sequence complementary to a first target site sequence of a target DNA and a second nucleotide sequence complementary to a second target site sequence on the opposite strand of the target DNA to the first target site sequence, and wherein the bridge RNA is capable of forming a complex with the IS110 family transposase.

[0012] In some embodiments, the bridge RNA further comprises a third nucleotide sequence complementary to a first donor site sequence of the donor DNA and a fourth nucleotide sequence complementary to a second donor site sequence on the opposite strand of the donor DNA to the first donor site sequence, hi some embodiments, the third nucleotide sequence and the fourth nucleotide sequence are on a second internal loop.

[0013] In some embodiments, the IS110 family transposase comprises a RuvC-like DEDD catalytic domain and a transposase domain. In some embodiments, the IS110 family transposase further comprises a linker domain between the RuvC-like DEDD catalytic domain and the transposase domain. In some embodiments, the linker domain comprises a coiled-coil linker domain. In some embodiments, the RuvC-like DEDD catalytic domain comprises an amino acid sequence at least 50% identical to a RuvC-like DEDD catalytic domain sequence of SEQ ID NOs: 10176-10523, 10524-20350, or 40357-516430. In some embodiments, the IS110 family transposase comprises a RuvC-like DEDD catalytic domain that forms a tertiary structure similar to the RuvC-like DEDD catalytic domain of IS621. In some embodiments, the RuvC-like DEDD catalytic domain of the IS110 family transposase comprises a tertiary structure similar to the tertiary structure of the RuvC-like DEDD catalytic domain of IS621 if the template modeling score (TM score) of the RuvC-like DEDD catalytic domain of the IS110 family transposase is 0.5 or greater. In some embodiments, the RuvC-like DEDD catalytic domain comprises an amino acid sequence that is at least 15% identical to the RuvC-like DEDD catalytic domain sequence of SEQ ID NOs: 10176-10523, 10524-20350, or 40357-516430.

[0014] In some embodiments, the transposase domain comprises an amino acid sequence at least 50% identical to the transposase domain sequence of SEQ ID NOs: 10176-10523, 10524-20350, or 40357-516430. In some embodiments, the IS110 family transposase comprises a transposase domain that forms a tertiary structure similar to the transposase domain of IS621. In some embodiments, if the template modeling score (TM score) of the transposase domain of the IS110 family transposase is 0.5 or greater, the IS110 family transposase domain comprises a tertiary structure similar to the tertiary structure of the transposase domain of IS621. In some embodiments, the transposase domain comprises an amino acid sequence at least 15% identical to the transposase domain sequence of SEQ ID NOs: 10176-10523, 10524-20350, or 40357-516430.

[0015] In some embodiments, the IS110 family transposase comprises an amino acid sequence at least 50% identical to SEQ ID NOs: 10176-10523, 10524-20350, or 40357-516430. In some embodiments, the IS110 family transposase further comprises a tertiary structure similar to that of IS621. In some embodiments, if the template modeling score (TM score) of the transposase is 0.5 or greater, the IS110 family transposase comprises a tertiary structure similar to that of IS621. In some embodiments, the transposase domain comprises an amino acid sequence at least 15% identical to the transposase domain sequence of SEQ ID NOs: 10176-10523, 10524-20350, or 40357-516430.

[0016] In some embodiments, the IS110 family transposase is IS110 group transposase. In some embodiments, the IS110 family transposase is an IS1111 group transposase. In some embodiments, the IS1111 group transposase is IS1111A or IS1111_229727. In some embodiments, the IS110 group transposase is IS621, ISPal 1, IsPa29, ISMmg1, ISPfl1, ISMae40, ISstma6, ISAzs32, ISMex9, ISCARP28, ISAar16, ISCps7, ISPpu9, ISRel9, ISEsa2, ISMma5, IS900, or ISHne5. In some embodiments, the IS110 group transposase comprises an amino acid sequence at least 50% identical to IS621 (SEQ ID NO: 10176).

[0017] In some embodiments, the bridge RNA comprises at least two stem-loop structures, including a first stem-loop and a second stem-loop, wherein the first stem-loop is 5' to the second stem-loop, the first stem-loop comprises a target-binding loop, and the second stem-loop comprises a donor-binding loop. In some embodiments, the bridge RNA further comprises a third stem-loop structure 5' to the first stem-loop. In some embodiments, the stem of the first stem-loop is 5-35 nucleotides long, the loop is 3-10 nucleotides long, and the target-binding loop is 5-20 nucleotides long, and the stem of the second stem-loop is 5-35 nucleotides long, the loop is 3-10 nucleotides long, and the donor-binding loop is 5-20 nucleotides long. In some embodiments, the stem of the second stem-loop structure comprises 1-4 loops or bubbles, each 1-10 nucleotides long.

[0018] In some embodiments, the bridge RNA comprises a nucleotide sequence comprising any of the 5' to 3' sequences shown in Figure 19, where "n" represents any nucleotide, "R" represents an A or G nucleotide, and "Y" represents a C or U nucleotide. In some embodiments, the bridge RNA comprises a 5' to 3' secondary structure shown in the first, second, third, or fourth column of the secondary structure of the sequence shown in Figure 19, where matching brackets "(" and ")" indicate base-paired nucleotides and "." indicates unpaired bases.

[0019] In some embodiments, the bridge RNA comprises a stem-loop structure as shown in FIG. 2D, FIG. 11B, or FIG.

[0020] In some embodiments, the target binding loop of the bridge RNA comprises: a left target guide (LTG) comprising, in a 5' to 3' direction, a nucleotide sequence complementary to a first strand of a target site sequence, wherein the 3' end of the LTG is complementary to at least one nucleotide of a core sequence on the first strand of the target site sequence; and a right target guide (RTG) comprising, in a 5' to 3' direction, a nucleotide sequence reverse-complementary to a strand opposite the first strand of the target site sequence, wherein the 3' end of the RTG is reverse-complementary to at least one nucleotide of a core sequence on the strand opposite the first strand of the target site sequence; is a polynucleotide sequence, and / or the donor binding loop of the bridge RNA comprises a left donor guide (LDG) comprising, in a 5' to 3' direction, a nucleotide sequence complementary to a first strand of a donor site sequence, wherein the 3' end of the LDG is complementary to at least one nucleotide of a core sequence on the first strand of the donor site sequence, and a right donor guide (RDG) comprising, in a 5' to 3' direction, a nucleotide sequence that is reverse complementary to a strand opposite the first strand of the donor site sequence, wherein the 3' end of the RDG is reverse complementary to at least one nucleotide of a core sequence on the strand opposite the first strand of the donor site sequence. RDG), wherein the donor site sequence is a polynucleotide sequence, and the core sequence of the donor site sequence is identical to the core sequence of the target site sequence.

[0021] In some embodiments, the target binding loop of the bridge RNA comprises: a left targeting guide (LTG) comprising, in a 5' to 3' direction, a nucleotide sequence that is reverse-complementary to a first strand of a target site sequence opposite the first strand, wherein the 3' end of the LTG is complementary to at least one nucleotide of a core sequence on the first strand of the target site sequence; and a right targeting guide (RTG) comprising, in a 5' to 3' direction, a nucleotide sequence that is complementary to the first strand of the target site sequence, wherein the 3' end of the RTG is complementary to at least one nucleotide of a core sequence on the first strand of the target site sequence, wherein the target site sequence is a polynucleotide sequence; and / or The loop comprises: a left donor guide (LDG) comprising, in a 5' to 3' direction, a nucleotide sequence complementary to a first strand of a donor site sequence, wherein the 3' end of the LDG is complementary to at least one nucleotide of a core sequence on the first strand of the donor site sequence; and a right donor guide (RDG) comprising, in a 5' to 3' direction, a nucleotide sequence that is the reverse complement of a strand opposite to the first strand of the donor site sequence, wherein the 3' end of the RDG is the reverse complement of at least one nucleotide of a core sequence on the strand opposite to the first strand of the donor site sequence; wherein the donor site sequence is a polynucleotide sequence, and the core sequence of the donor site sequence is identical to the core sequence of the target site sequence.

[0022] In some embodiments, the target site sequence is the sequence X1X2X3X4X5X6X7X8X9X 10 X 11 X 12 X 13 X 14 (where X is any nucleotide, X8X9 is the core, and X 12 , X 13 , and X 14wherein one or more of the following are optionally part of the target site sequence. In some embodiments, the bridge RNA comprises a 5' to 3' sequence of LTG X1X2X3X4X5X6X7X8 and a 5' to 3' sequence of Y 14 Y 13 Y 12 Y 11 Y 10 Y9Y8 RTG, where Y is the complementary nucleotide to X, and Y 14 , Y 13 , and Y 12 is optionally part of an RTG, or the bridge RNA comprises, in the 5' to 3' direction, an LTG of X1X2X3X4X5X6X7X8X9 and, in the 5' to 3' direction, a Y 14 Y 13 Y 12 Y 11 Y 10 Y RTG, where Y is the complementary nucleotide to X, and Y 14 , Y 13 , and Y 12 In some embodiments, one or more of the donor site sequences are selected from the group consisting of the sequence STIR-X, n1 -X1X2X3X4X5X6X7X8X9X 10 X 11 X 12 X 13 X 14 -X n2 -STIR (where X is any nucleotide) 12 , X 13 , and X 14 wherein one or more of the following are optionally part of the donor site sequence; STIR is optional but, if present, is a subterminal inverted repeat comprising 2-20 nucleotides; X8X9 is the core; and n1 and n2 can independently be 0-10. In some embodiments, the bridge RNA comprises a LDG of X1X2X3X4X5X6X7X8 in the 5' to 3' direction and a Y in the 5' to 3' direction. 14 Y 13 Y 12 Y 11 Y 10Y9Y8 RDG, where Y is the complementary nucleotide to X, and Y 14 , Y 13 , and Y 12 One or more of the following may optionally be part of an RDG, or the bridge RNA may comprise an LDG of X1X2X3X4X5X6X7X8X9 in the 5' to 3' direction and a Y 14 Y 13 Y 12 Y 11 Y 10 Y9 RDG, where Y is the complementary nucleotide to X, and Y 14 , Y 13 , and Y 12 In some embodiments, one or more of the STIR sequences are optionally part of an RDG. In some embodiments, the STIR sequence, if present, comprises a G / T-rich nucleotide sequence. In some embodiments, the 5' STIR sequence, if present, comprises a G / T-rich nucleotide sequence.

[0023] In some embodiments, the target site sequence is located on genomic DNA, linear dsDNA, dsDNA plasmid, ssDNA, or RNA. In some embodiments, the donor site sequence is located on genomic DNA, linear dsDNA, dsDNA plasmid, ssDNA, or RNA. In some embodiments, the dsDNA plasmid further comprises a polynucleotide sequence for insertion into the donor site sequence. In some embodiments, the dsDNA plasmid further comprises a polynucleotide sequence for insertion into the target site sequence. In some embodiments, the target site sequence and donor site sequence on the genomic DNA are located on the same DNA strand. In some embodiments, the target site sequence and donor site sequence on the genomic DNA are located on different chromosomes.

[0024] In some embodiments, the bridge RNA is a split bridge RNA. In some embodiments, the bridge RNA comprises a first RNA molecule comprising a first portion of the bridge RNA and a second RNA molecule comprising a second portion of the bridge RNA. In some embodiments, the first portion of the bridge RNA comprises an internal loop comprising the first nucleotide sequence complementary to the first target site sequence of the target DNA, and the second portion of the bridge RNA comprises a second internal loop comprising a third nucleotide sequence complementary to a first donor site sequence of the donor DNA and a fourth nucleotide sequence complementary to a second donor site sequence. In some embodiments, the first portion of the bridge RNA is encoded on a nucleic acid and operably linked to a first promoter, and the second portion of the bridge RNA is encoded on the same or a different nucleic acid and operably linked to a second promoter. In some embodiments, the first and second portions of the bridge RNA are encoded on nucleic acids and further comprise one or more ribozyme sites that cleave the expressed bridge RNA into the first and second RNA molecules.

[0025] In some embodiments, any of the LTG, RTG, LDG, and / or RDG of the bridge RNA is complementary to more nucleotides of the respective target site sequence or donor site sequence than the number of complementary nucleotides in the corresponding naturally occurring bridge RNA. In some embodiments, the total number of nucleotides comprising the target binding loop and / or donor binding loop of the bridge RNA is the same as the corresponding naturally occurring bridge RNA. In some embodiments, the total number of nucleotides comprising the target binding loop and / or donor binding loop of the bridge RNA is increased compared to the corresponding naturally occurring bridge RNA. In some embodiments, any of the LTG, RTG, LDG, and / or RDG of the bridge RNA is not complementary to nucleotides of the core sequence in the respective target site sequence or donor site sequence, but the number of complementary nucleotides in the respective target site sequence or donor site sequence is the same as in the corresponding naturally occurring bridge RNA. In some embodiments, the target site sequence is a sequence X -1 X1X2X3X4X5X6X7X8X9X 10 X 11 X 12 X 13 X 14 (where X is any nucleotide, X8X9 is the core, and X 12 , X 13 , and X 14 one or more of which are optionally part of the target site sequence), and the bridge RNA comprises X in the 5' to 3' direction -1 X1X2X3X4X5X6X7 LTG and Y in the 5' to 3' direction 14 Y 13 Y 12 Y 11 Y 10 Y9Y8 RTG, where Y is the complementary nucleotide to X, and Y 14 , Y 13 , and Y 12 In some embodiments, one or more of the donor site sequences are selected from the group consisting of the sequence STIR-X, n1 -X1X2X3X4X5X6X7X8X9X10 X 11 X 12 X 13 X 14 -X n2 -STIR (where X is any nucleotide) 12 , X 13 , and X 14 one or more of which are optionally part of the donor site sequence; STIR is optional, but if present is a subterminal inverted repeat sequence containing 2 to 20 nucleotides; X8X9 is the core; n1 and n2 can independently be 1 to 10; and the bridge RNA comprises X in the 5' to 3' direction. -1 X1X2X3X4X5X6X7 and Y on RDG in the direction from 5' to 3' 14 Y 13 Y 12 Y 11 Y 10 Y9Y 8 LDG, where Y is the complementary nucleotide to X, and Y 14 , Y 13 and Y12 are optionally part of RDG. In some embodiments, one or more of the LTG, RTG, LDG, and / or RDG nucleotides of the bridge RNA base pair with nucleotides of the respective target site sequence or donor site sequence via non-canonical base pairing.

[0026] In certain aspects, the invention provides vectors comprising any of the nucleic acids of the nucleic acid editing systems of the invention. In certain aspects, the invention provides host cells comprising any of the vectors of the invention. In some embodiments, any of the nucleic acids of the nucleic acid editing systems further comprises an inducible promoter.

[0027] In certain aspects, the present invention provides a method for integrating a DNA molecule of interest into a sequence-specific site in a DNA of interest in a cell, said method comprising introducing a nucleic acid editing system of the present invention into the cell.

[0028] In some embodiments, the cell is a mammalian cell, hi some embodiments, the cell is a human cell.

[0029] In some embodiments, the DNA of interest of the cell comprises a donor site sequence, and the DNA molecule of interest is part of the nucleic acid of the nucleic acid editing system, which comprises a target site sequence. In some embodiments, the DNA of interest of the cell comprises a target site sequence, and the DNA molecule of interest is part of the nucleic acid of the nucleic acid editing system, which comprises a donor site sequence. In some embodiments, the DNA of interest of the cell further comprises a second donor site sequence, the DNA molecule of interest further comprises a second target site sequence, and the nucleic acid editing system comprises a second bridge RNA that targets the second donor site sequence and the second target site sequence. In some embodiments, the DNA of interest of the cell comprises a second target site sequence, the DNA molecule of interest further comprises a second donor site sequence, and the nucleic acid editing system comprises a second bridge RNA that targets the second donor site sequence and the second target site sequence. In some embodiments, the sequence of the bridge RNA is engineered prior to introduction of the nucleic acid editing system to bind to the donor site sequence and the target site sequence. In some embodiments, the DNA of interest in the cell is the genome of the cell, hi some embodiments, the DNA of interest in the cell is a plasmid.

[0030] In certain aspects, the present invention provides a method for inverting a DNA sequence of a DNA of interest in a cell, comprising introducing a nucleic acid editing system of the present invention into the cell, wherein a target site sequence and a donor site sequence are present on the same DNA molecule of interest, and the LD of the donor site sequence and the RT of the target site sequence are on the same DNA strand.

[0031] In some embodiments, the DNA of interest in the cell is the genome of the cell. In some embodiments, the sequence of the bridge RNA is engineered to bind to the donor site sequence and the target site sequence prior to introduction of the nucleic acid editing system.

[0032] In certain aspects, the present invention provides a method for excising a DNA sequence in a DNA of interest in a cell, the method comprising introducing a nucleic acid editing system of the present invention into the cell, wherein a target site sequence and a donor site sequence are present on the same DNA molecule of interest, and the LD of the donor site sequence and the LT of the target site sequence are on the same DNA strand.

[0033] In some embodiments, the DNA of interest in the cell is the genome of the cell. In some embodiments, the sequence of the bridge RNA is determined by: It is engineered to bind to the donor site sequence and the target site sequence.

[0034] In certain aspects, the present invention provides a method for translocating a DNA sequence between two linear DNA molecules of interest, comprising introducing a nucleic acid editing system of the present invention into a cell, wherein a donor site sequence is present on a first linear DNA molecule and a target site sequence is present on a second linear DNA molecule.

[0035] In some embodiments, the linear DNA molecule of interest in the cell is a chromosome of the cell. In some embodiments, the sequence of the bridge RNA is engineered to bind to the donor site sequence and the target site sequence prior to introduction of the nucleic acid editing system.

[0036] Other embodiments of the present invention are further described in the following sections of this application, including the Brief Description of the Drawings, Detailed Description, and Examples, and in the claims. Still other objects and advantages of the present invention will be apparent to those skilled in the art from the disclosure herein, which is intended to be illustrative and not limiting. Accordingly, other embodiments will be apparent to those skilled in the art without departing from the spirit and scope of the present invention.

[0037] The patent or application file contains at least one drawing originally executed in color. To conform to PCT patent application requirements, many of the figures presented herein are black-and-white representations of images originally executed in color. [Brief explanation of the drawings]

[0038] [Figure 1A]Figures 1A-E show general features of IS110 insertion sequence elements. (A) Sequence characteristics of two groups of IS110 elements. The IS110 group is characterized by a longer left non-coding end (LE) and a shorter right non-coding end (RE). The IS1111 group is characterized by a shorter LE and a longer RE. In both groups, a core sequence motif (1-5 nt) is found at both ends of the element. IS110 was previously thought to lack a subterminal inverted repeat (STIR), while IS1111 was known to have a 6-12 nt subterminal inverted repeat. However, as described herein, most IS110 elements also have a short STIR (see Figures 10A-B). IS110 elements are typically 1000-2000 nt in length. This figure discloses SEQ ID NOs: 795157-795160. (B) Diagram of the IS110 transposase domains. IS110 transposases are typically 300-500 aa in length. The RuvC-like domain (DEDD_Tnp_IS110 by Pfam) contains the canonical DEDD catalytic motif. The IS110 Tnp domain (Transposase_20 by Pfam) contains the catalytic serine. (C) Illustrated life cycle of IS110 elements. Once integrated into the genome, IS110 elements excise themselves from the genome, resulting in intact repair of the genomic DNA target site and the formation of a circular IS110 element. In the circular form, the RE becomes adjacent to the LE. Insertion can occur into the same dsDNA target site or a new target site. The inserted linear IS110 element consists of a left non-coding end (LE), a coding sequence for the transposase (Tpase), and a right non-coding end (RE). The inserted IS110 element is flanked on the left end by a left flank (LF) (box at the left end) containing the left target (LT) sequence, and on the right end by a right flank (RF) (box at the right end) containing the right target (RT) sequence. Many IS110 elements contain identical "core" sequences (diamonds) between the LF and LE, and between the RE and RF, although not all IS110 elements utilize these sequences.The IS110 element excises itself, resulting in a pre-insertion ("target") site with LF-core (if present)-RF and a circular element with RE-core (if present)-LE-Tpase. The ligation of the RE-LE junction forms a "donor" site sequence as a subsequence of the RE-LE junction, which includes other core sequences, if present, found in the integrated element. The donor site sequence may also contain a subterminal inverted repeat (STIR) sequence, indicated by a triangle, although STIR is not essential for IS110 recombinase activity. The RE-LE ligation also forms a promoter that can drive expression of a bridge RNA from the RE or LE. It may also drive expression of a transposase. The circular form of the element can reinsert into the target site from which it was excised or into any other polynucleotide containing a target site sequence, and the bridge RNA encoded within the LE or RE recognizes the donor site sequence and / or the target site sequence to mediate transposition. This figure discloses SEQ ID NOS: 795161-795164. (D) Phylogenetic tree of IS110 transposases. IS110 has several clades, but is generally distinguished by the IS110 group and the IS1111 group. Host kingdoms and phyla are indicated to indicate diverse origins. Notable locations of IS110 transposases are highlighted on the tree. (E) Comparison of end lengths of IS110 groups. IS110 typically has a longer LE than RE, while IS1111 typically has a longer RE than LE. [Figure 1B] Same as above [Figure 1C] Same as above [Figure 1D] Same as above [Figure 1E] Same as above

[0039] [Figure 2A]Figures 2A-D show the identification of a bridge RNA from the model IS110 IS621. (A) RNA sequencing of the IS110 non-coding end (sequence number 795165). A plasmid encoding the RE-LE sequence was delivered to E. coli, and RNA was extracted and sequenced. The boundaries of the RNA encoded within the LE are defined across six IS621 orthologs. (B) Demonstration of bridge RNA binding to IS621 transposase. The IS621 RNA from part A was purified and exposed to various concentrations of IS621 transposase. Microscale thermophoresis was used to measure the rate of bridge RNA binding to the transposase. A scrambled RNA with mismatched bases and the reverse complement of the bridge RNA served as negative controls. (C) Determination of IS621 bridge RNA structure. Hundreds of related IS621 LEs were aligned, and RNA structures were predicted for each. The dominant structures at each position in the alignment were calculated and plotted. The structures were characterized as 5' stem, 3' stem, hairpin, other, or gap. Between the LE and CDS initiations, a structure characterized by a consensus bridge RNA structure and an accessory structure at the 5' end emerged. (D) Illustration of the IS621 bridge RNA structure: nnnnnnnnYYnRRnnnnYYYYnnnYnnnnRRnnnnYYGGAYGCCGYnYYnRnCCUnnRRYnnnARYYYGYnnYGUAGAUnnnYGCRnCRRnYRYYnnnnnnnnYnnGYnnnRRRYCGRACnGnAUCnYnGGCYGGYnnnYCGRnARYCYGCAUYACAAGUnGRUnRCRYRAnnnn (SEQ ID NO: 795166). The information in C is rendered here using the software R2R. The target-binding loop and donor-binding loop are labeled. The accessory structures of this particular bridge RNA are also labeled.The secondary structure of the IS621 bridge RNA structure in Figure 2D is represented in "dot bracket" notation as follows: ...((((((((((....)))))))))).((((((((((.(((............(((((....)))))..............))))))))))))).......(.((((....(((.............(((((((....)))))...).............))))))).)..., where matching brackets "(" and ")" indicate base pairs and unpaired bases are indicated by dots ("."). [Figure 2B] Same as above [Figure 2C-1] Same as above [Figure 2C-2] Same as above [Figure 2D] Same as above

[0040] [Figure 3A-1]Figures 3A-D show the prediction and validation of the DNA recognition mechanism by the bridge RNA. (A) Illustrated covariation analysis approach. The boundaries of the IS110 element are identified using comparative genomics. The non-coding ends are ligated to identify the donor site, predicting a circular IS110 structure, and the pre-integration target site is extracted. The bridge RNA sequence is predicted from the non-coding end. The alignment of the target site (or donor site) is compared with the structural alignment of the bridge RNA sequence to identify covarying positions (sequence numbers 795167-795168, respectively, in order of appearance). (B) Covariation between the IS621 bridge RNA and its target and donor. The covariation scores calculated by the software CCMpred were normalized and plotted for each target and donor position along the left end. Subsequences of the bridge RNA covary with the target and donor. Covarying sequences are observed to be complementary or reverse-complementary to the donor and target. These covariant regions in the bridge RNA were then examined for evidence of base pairing with the target and donor sequences to identify programmable guide sequences. This figure discloses SEQ ID NOs: 795172, 795164, and 795169-795171, respectively, in order of appearance. (C) Model of the IS621 bridge RNA with the target and donor sequences. A representation of the R2R structure from Figure 2D is shown, along with the positions of the LTG, RTG, LDG, and RDG within the target and donor binding loops. The LT, RT, LD, and RD are shown with their relative positions to the core sequence found in both the target and donor. The subterminal inverted repeats are also shown on the donor. The target binding loop of the bridge RNA contains the target site sequences, i.e., the left target guide (LTG) and right target guide (RTG), specific for the left target (LT) and right target (RT) sequences, respectively. In the case of IS110 family transposases that use a core sequence, both the LTG and RTG may contain base-pairing specificity for at least one base of the core dimer sequence (dashed line). The donor-binding loop of the bridge RNA contains the left donor guide (LDG) and right donor guide (RDG) sequences specific for the donor site sequences, i.e., the left donor (LD) and right donor (RD) sequences, respectively.In the case of IS110 family transposases that use a core sequence, both the LDG and RDG may contain base-pairing specificity for at least one base of the core sequence. The donor site may also encode a subterminal inverted repeat (STIR) that directly interacts with the transposase protein and may therefore be required for other parts of the IS110 life cycle, such as the transposition reaction or cut and paste. In some embodiments, additional bridge RNA nucleotides outside of these described guide sequences may play a role in programming specificity for different target and donor sequences. This figure discloses SEQ ID NOs: 795172 and 795164, respectively, in order of appearance. (D) Demonstration of sequence-specific binding of the target and donor. IS621 transposase and bridge RNA ribonucleoprotein were exposed to WT target and donor sequences, as well as a scrambled DNA sequence matching at position 0. The binding affinity of the WT bridge RNA for the WT donor and target is shown. [Figure 3A-2] Same as above [Figure 3B-1] Same as above [Figure 3B-2] Same as above [Figure 3C] Same as above [Figure 3D] Same as above

[0041] [Figure 4-1]Figure 4 shows a diagram of how a bridge RNA can be reprogrammed to recognize new targets and donors. A diagram of a bridge RNA programmed to recognize new targets and donors is shown. The WT bridge RNA is shown first with the WT target and donor sequences. The target and / or donor loops are modified in each of the following examples. The bridge RNA sequence is shown with the indicated LTG, RTG, LDG, and RDG sequences, which can base-pair with the illustrated target and donor site sequences. These subsequences can be reprogrammed to bind any desired target or donor sequence. In all five examples, the core nucleotides of the target and donor match each other. In the final example, STIR is modified to any sequence, as it is not strictly required for transposition function. This figure discloses SEQ ID NOs: 795169, 795173-795174, 795170-795171, 795175-795178, 795171, 795179, 795176, 795180, 795178, 795180-795184, 795183, 795185-795186, 795187-795188, and 795187, respectively, in order of appearance. [Figure 4-2] Same as above [Figure 4-3] Same as above [Figure 4-4] Same as above

[0042] [Figure 5A]Figures 5A-D show in-cell demonstration of transposition and target reprogramming. (A) GFP reporter assay for transposition. The pDonor plasmid encodes an inactive GFP gene flanked by the RE-LE of IS621, which expresses a bridge RNA. pDonor is cotransformed into E. coli with pTarget, which encodes the WT target and IS621 transposase, both flanked by promoters. Integration of the donor into the target activates GFP expression. (B) Demonstration of GFP reporter assay function. E. coli cells are measured for GFP expression in the FITC channel using flow cytometry. GFP expression is observed only when the transposase is WT, using the assay in (A); inactivation of conserved residues in either the RuvC-like domain or the Tnp domain abolishes transposition. (C) Diagram of reprogramming of the bridge RNA to recognize a new target. The target loop of the bridge RNA is modified to recognize a new target. The target is also modified to match the target loop when used in the reporter assay in A. The graph demonstrates specific retargeting of transposition to a new target. Plasmids expressing seven bridge RNAs with unique target loops were paired with matching or wild-type targets. Transposition was observed only when a matching target was provided, as measured by flow cytometry for GFP expression. The figure discloses SEQ ID NOs: 795189, 795190, 795176, 795191-795194, and 795173, respectively, in order of appearance. (D) Isolation of bridge RNA from RE-LE. The bridge RNA can be isolated and expressed from a different promoter to achieve higher transposition rates. The donor sequence can be reduced from 298 bp to at least 22 bp without affecting transposition efficiency. Shortening to 11 bp (removal of STIR) reduces transposition efficiency to near background. When the bridge RNA is present with the 11 bp donor, some integration is observed. In systems lacking the bridge RNA, integration never occurs. This figure discloses SEQ ID NOs: 795195 to 795196 in the order in which they appear. [Figure 5B] Same as above [Figure 5C]Same as above [Figure 5D] Same as above

[0043] [Figure 6A] Figures 6A-E show the IS621 bridge RNA target / target loop mismatch tolerance and reprogramming screen. (A) Schematic of the antibiotic resistance reporter design. A minimal donor (22 bp) is encoded on a plasmid adjacent to a kanamycin resistance gene. A second plasmid encodes the target, bridge RNA, and transposase. The target is linked to the bridge RNA using a barcode. Recombination between the donor and target plasmids results in survival of the E. coli, and functional bridge RNA target loop-target pairs are recorded using next-generation sequencing. (B) Schematic of the target specificity screen design. The target (SEQ ID NOs: 795197-795205, respectively, in order of appearance) and target loop are varied, except for the target core and the LTG and RTG subsequences that bind to the core. The donor loop (and donor) are kept constant. Target-target loop pairs are designed to assay for single mismatches, double mismatches, and total mismatches. Targets in the screen are selected to reduce the number of off-targets in the E. coli genome. (C) Abundance of target and target loop pairs. Abundance is measured by barcode counts per million reads. Zero-mismatch target / target loop pairs are generally abundant, while increasing the number of mismatches decreases abundance. (D) Sequence logos of top quintile targets. The relative enrichment of nucleotides at each position in the target is shown for the zero-mismatch target / target loop pairs in the top quintile of 6364 target / target loop pairs. (E) Single mismatch tolerance by position. For the top quintile of target sets, the relative enrichment of nucleotides is shown whether the target loop is mismatched to each position in the target (sequence number 795173). The best-performing zero-mismatch pair in each target set was used to represent the set, and the top quintile of target sets is shown. [Figure 6B] Same as above [Figure 6C] Same as above [Figure 6D]Same as above [Figure 6E] Same as above

[0044] [Figure 7A]Figures 7A-H show the mismatch tolerance and reprogramming screen for the IS621 bridge RNA donor / donor loop. (A) Schematic of the antibiotic resistance reporter design. The full-length donor is encoded on a plasmid adjacent to a kanamycin resistance gene. A constitutive promoter from the WT system expresses the bridge RNA. A unique molecular identifier (UMI) identifies the donor / donor loop pair. A second plasmid encodes the target and transposase. Recombination between the donor and target plasmids results in survival of the E. coli, and functional bridge RNA donor loop and donor pairs are recorded using next-generation sequencing. (B) Schematic of the donor specificity and reprogramming design. The donor (sequence numbers 795206-795212, respectively, in order of appearance) and donor loop are varied, except for the core and the LDG and RDG subsequences that bind to the core. The target loop (and target) are kept constant; however, the target is a non-WT sequence not found in the E. coli genome. Target and target loop pairs are designed to assay single and double mismatches for the WT donor sequence. Five thousand random perfectly matched donor and donor loops are also assayed. (C) UMI abundance for donor / donor loop pairs that differ by 1 nt from WT. Counts are plotted for pairs with zero or one mismatch between the donor and donor loop. The UMI abundance of the WT donor paired with the WT donor loop is depicted by a red dashed line. (D) UMI abundance for donor / donor loop pairs that differ by 2 nt from WT. Counts are plotted for pairs with zero or two mismatches between the donor and loop. The UMI abundance of the WT donor paired with the WT donor loop is depicted by a red dashed line. (E) UMI abundance for donor / donor loop pairs with zero mismatches. CPM values ​​for donor / donor loop pairs are binned by the number of nucleotide differences between the reprogrammed donor and the WT donor. The CPM of the WT donor paired with the WT donor loop is shown with a dashed red line. (F) Single mismatch tolerance by position in the donor (sequence number 795196). The relative enrichment of nucleotides for the top quintile of the donor set is shown along with all 4 x 4 = 16 mismatch combinations tested at each position in the donor.(G) Sequence logos of the top quintile donors. The relative enrichment of nucleotides at each target position is shown for the top quintile of 5,000 donor / donor loop pairs with zero mismatches. (H) Demonstration of specific retargeting of transfer to a new donor. Plasmids expressing five bridge RNAs (sequence numbers 795196 and 795213-795221, respectively, in order of appearance) with unique donor loops were matched with the cognate donor or the WT donor. Transfer was observed only when a matched donor was provided, as measured by flow cytometry for GFP expression with FITC. Results were generated using a 22-bp donor using the approach in Figure 5D. [Figure 7B] Same as above [Figure 7C] Same as above [Figure 7D] Same as above [Figure 7E] Same as above [Figure 7F] Same as above [Figure 7G] Same as above [Figure 7H] Same as above

[0045] [Figure 8A]Figures 8A-C show diagrams and demonstrations of DNA reorganization by IS621 transposase. (A) Diagram of a GFP reporter assay for DNA insertion. Plasmids encoding donor and GFP coding sequences and a target plasmid flanked by promoters are introduced into E. coli. Coexpression of bridge RNAs encoding target and donor loops matching the donated target and donor results in efficient insertion in E. coli. This figure discloses SEQ ID NOS: 795164 and 795222-795224, respectively, in order of appearance. (B) Diagram of a GFP reporter assay for DNA excision recombination. Plasmids encoding a promoter flanking the donor and a target preceded by a terminator and followed by a GFP coding sequence are delivered to E. coli. Coexpression of bridge RNAs encoding target and donor loops matching the donated target and donor results in efficient excision recombination in E. coli, and removal of the intervening sequence encoding the terminator allows GFP expression. This reaction results in one DNA molecule becoming two DNA molecules. This figure discloses SEQ ID NOS: 795164, 795225, 795223, and 795226, respectively, in order of appearance. (C) Diagram of GFP reporter assay for DNA inversion. A plasmid encoding a promoter adjacent to a donor and a target preceded by a terminator and GFP coding sequence is delivered to E. coli. Coexpression of a bridge RNA encoding a target and donor loop matching the donor and target results in efficient inversion in E. coli, and inversion of the sequence between the donor and target allows GFP expression. This figure discloses SEQ ID NOS: 795164, 795227, 795226, and 795228, respectively, in order of appearance. (A-C) Insertion efficiency is measured by the percentage of cells expressing GFP as determined by flow cytometry. Excision recombination efficiency is measured by the percentage of cells expressing GFP as determined by flow cytometry. Inversion efficiency is measured by the percentage of cells expressing GFP as determined by flow cytometry. [Figure 8B] Same as above [Figure 8C] Same as above

[0046] [Figure 9A] Figures 9A-B show a diagram and demonstration of DNA insertion into the E. coli genome using IS621 transposase. (A) Illustrated genome integration assay. Due to kanamycin resistance encoded on the donor plasmid, an E. coli strain containing the donor plasmid is generated that can grow under kanamycin selection below 37°C. At 37°C, the donor plasmid cannot replicate. The pHelper plasmid, encoding the IS621 transposase and a bridge RNA that recognizes the donor on the plasmid and a target within the genome, is delivered to E. coli. Growth at 37°C results in the selection of E. coli that have integrated the plasmid into their genome, which must survive. (B) Integration profile using bridge RNAs targeting four sites within the genome for integration. Targets are ranked by the relative abundance of the integration location from nanopore sequencing data. Integration sites are color-coded by the number of differences between the observed and expected target sites. This figure discloses SEQ ID NOs: 795229-795235, respectively, in order of appearance. [Figure 9B] Same as above

[0047] [Figure 10A] Figures 10A-B show the identification of subterminal inverted repeats in IS110 elements of the IS110 group. (A) Diagram of the approach for identifying subterminal inverted repeats. The boundaries of IS110 elements are identified using comparative genomics and BLAST. Non-coding ends are linked and aligned to predict circular conformation. Covarying sequences are compared across donors up to 25 bp in each direction from the core. (B) Sequence covariation within donors identifies short subterminal inverted repeats. Covariation scores are plotted for each position in the donor for covariation with itself. [Figure 10B] Same as above

[0048] [Figure 11A-1]Figures 11A-C show the prediction and validation of bridge RNAs expressed from IS1111 REs. (A) Determination of IS1111_229727 bridge RNA structure. Hundreds of related IS1111_229727 REs were aligned, and RNA structures were predicted for each. The major structures at each position in the alignment were calculated and graphed. The structures were characterized as 5', 3' stem, hairpin, other, or gap. Between the predicted RE start and the element boundary, a structure characterized by a target-binding loop and a donor-binding loop appears. (B) Diagram of IS1111_229727 bridge RNA structure nnnnGnYYnYYGRnRGnGRCGYRGCCCGGYnnnnGnGYAAYCCnCGnnnnYRYnnnGnRRCnRnnnYnRAYYYnCGnnnYYAGAnnGnGGCAGGCnCnYnnGRCGGAnnYnnGYRnnGYGGUAnCCARYCCRCGRAURUCAGCnnGAUYnRCCGUCGnnnnnnRCYnGCYnCGYCnCnnYYRRnnRnYnnnnn (SEQ ID NO: 795236). The information in A is represented here using the software R2R. The secondary structure of the IS1111_229727 bridge RNA structure in Figure 11B is represented in "dot bracket" notation as follows: ...((((((((((((.(((((((((((..((((((((....(((((((((((((..........((((((((((((((.)))))..))))))))..........)))))...))))....(((((..........(((((((((((.)))))..))))))))))))))))))))))..., where matching brackets "(" and ")" indicate base pairs, and unpaired bases are indicated by dots ("."). (C) RNAseq validation of the IS1111_229727 bridge RNA. RNAseq coverage is illustrated above the RE of IS1111_229727. [Figure 11A-2] Same as above [Figure 11B] Same as above [Figure 11C] Same as above

[0049] [Figure 12A-1] Figures 12A-B show alignments of the RuvC and Tnp domains of various IS110 transposases. (A) Alignment of IS110 RuvC-like domains (SEQ ID NOS: 795237-795261, respectively, in order of appearance). The alignment is illustrated with conserved residues and regions. Residues are color-coded by amino acid chemical properties. (B) Alignment of IS110 Tnp domains (SEQ ID NOS: 795262-795286, respectively, in order of appearance). The alignment is illustrated with conserved residues and regions. Residues are color-coded by amino acid chemical properties. [Figure 12A-2] Same as above [Figure 12B] Same as above

[0050] [Figure 13-1] Figure 13 shows the various predicted bridge RNA structures associated with various IS110 transposases. The various bridge RNA consensus structures predicted across various IS110 transposases are shown. The procedure for generating each structure was the same as that used to generate the IS621 bridge RNA consensus structure. The RNA covariation models were clustered using a graph clustering approach, and consensus structures from 12 different clusters are shown. Each structure contains at least one loop similar to the target and / or donor loop. Significantly covarying base pairs are highlighted in gray boxes. The bridge RNA structures in Figure 13 are representations of sequence numbers 795287 to 795303, respectively, with gap positions excluded and excess unstructured bases trimmed. [Figure 13-2] Same as above [Figure 13-3] Same as above [Figure 13-4] Same as above [Figure 13-5] Same as above [Figure 13-6] Same as above [Figure 13-7] Same as above [Figure 13-8] Same as above [Figure 13-9] Same as above [Figure 13-10] Same as above [Figure 13-11] Same as above [Figure 13-12] Same as above [Figure 13-13] Same as above [Figure 13-14] Same as above [Figure 13-15] Same as above [Figure 13-16] Same as above [Figure 13-17] Same as above

[0051] [Figure 14A]Figures 14A-H show the tertiary structure alignment and analysis of IS110 transposase proteins. (A) Formula for template modeling score (TM score), where L is the length of the amino acid sequence of the target protein, and L is the number of residues that appear in both the template and target structures. d is the distance between the i-th residue pair in the template and target structures, and do(L) = 1.243√(L-15)-1.8 is the distance scale that normalizes the distance (Zhang and Skolnick 2004). Alternatively, the score can be normalized according to the length of the query protein or the average length of the two proteins. The TM score takes values ​​in (0, 1), and a cutoff of >0.5 is commonly used to identify proteins with homologous tertiary structures (Zhang and Skolnick 2004). (2005). (B) TM score distribution when the predicted IS110 structure is aligned to the AlphaFold structure of IS621. Each row shows the distribution of TM scores when normalized according to the length indicated on the right: the average of the two lengths, the length of IS621, or the length of the query protein. The dotted line indicates a TM score of 0.5, the minimum score threshold commonly used to identify homologous proteins. (C) Structural alignment of two distantly related IS110 proteins. IS621 is shown in green, and another predicted IS110 transposase structure is shown in cyan. Four different angles of the same structural alignment are shown. The two proteins are 18.1% identical at the amino acid level, but have a TM score of 0.805. (D) TM score distribution of the IS110 structure when clustered and aligned to the IS621 structure. Protein structures were clustered at 100%, 90%, and 50% identity, and a representative example from each cluster was chosen. TM scores normalized by the average length of the two sequences are shown. Each panel represents a different level of amino acid identity clustering. (E) TM scores of the RuvC and Tnp domains when aligned to the IS621 domain compared to the TM score of the complete protein.Domains were extracted using boundaries identified by the corresponding Pfam domains (DEDD_Tnp_IS110 and Transposase_20). These domain substructures were then aligned to the IS621 substructure using TM-align. TM scores are shown for the complete protein, the RuvC domain, and the Tnp domain. (F) IS630 transposase TM score versus IS110 transposase TM score. All IS110 and IS630 family transposase structures were aligned to the IS621 AlphaFold structure. TM scores normalized by the average length of the two sequences are shown. The IS630 family was chosen for comparison because it had a protein length distribution similar to that of IS110. (G) Schematic diagram showing the location of conserved residues within each protein structural domain and the estimated distances between them. The top panel shows the five conserved residues of the RuvC domain in a representative IS110 structure and a representative IS1111 structure. Residues are color-coded and labeled in red. Five positions are labeled P1–P5. The calculated distances between these residues are also shown; D1–D3 are color-coded and labeled in blue, purple, and green, respectively. The bottom panel shows the five conserved positions in the Tnp domain and the three calculated distances. Distances are relative to the alpha carbon of each residue. (H) Distances between conserved residues in the RuvC and Tnp domains of the AlphaFold structure of IS110. Distances were calculated as described in the previous paragraph. The distribution of distances in angstroms (Å) for each distance within each domain is shown here. See Figure 14G for distance reference. [Figure 14B] Same as above [Figure 14C] Same as above [Figure 14D] Same as above [Figure 14E] Same as above [Figure 14F] Same as above [Figure 14G] Same as above [Figure 14H] Same as above

[0052] [Figure 15-1] Figure 15 provides a sequence listing of IS110 elements (SEQ ID NOS: 1-348). Elements are represented as 5'-3' nucleotide sequences in typical FASTA format with additional formatting to indicate subsequences of interest. When available, annotations include: a dark gray highlight at the beginning of the sequence, indicating the core. The core is shown only once and, if annotated, is always at the 5' end. Light gray highlights always indicate the LE and RE flanking the CDS sequence. This is simply defined as the sequence that comes between the CDS and the end of the core or element. The CDS sequence is shown as a single-underlined, unhighlighted sequence. Predicted boundaries of the bridge RNA are indicated by lowercase nucleotides. If present, the guide sequence is shown in bold. If present, the four bolded subsequences represent LTG, RTG, LDG, and RDG, in that order. Additional IS110 elements are provided as SEQ ID NOS: 349-10175 in the attached sequence listing, which are incorporated herein by reference in their entirety. The sequence listing includes the start and stop positions of the core, LE, RE, CDS, and bridge RNA sequences as features of this sequence listing. [Figure 15-2] Same as above [Figure 15-3] Same as above [Figure 15-4] Same as above [Figure 15-5] Same as above [Figure 15-6] Same as above [Figure 15-7] Same as above [Figure 15-8] Same as above [Figure 15-9] Same as above [Figure 15-10] Same as above [Figure 15-11] Same as above [Figure 15-12] Same as above [Figure 15-13] Same as above [Figure 15-14] Same as above [Figure 15-15] Same as above [Figure 15-16] Same as above [Figure 15-17] Same as above [Figure 15-18] Same as above [Figure 15-19] Same as above [Figure 15-20] Same as above [Figure 15-21] Same as above [Figure 15-22] Same as above [Figure 15-23] Same as above [Figure 15-24] Same as above [Figure 15-25] Same as above [Figure 15-26] Same as above [Figure 15-27] Same as above [Figure 15-28] Same as above [Figure 15-29] Same as above [Figure 15-30] Same as above [Figure 15-31] Same as above [Figure 15-32] Same as above [Figure 15-33] Same as above [Figure 15-34] Same as above [Figure 15-35] Same as above [Figure 15-36] Same as above [Figure 15-37] Same as above [Figure 15-38] Same as above [Figure 15-39] Same as above [Figure 15-40] Same as above [Figure 15-41] Same as above [Figure 15-42] Same as above [Figure 15-43] Same as above [Figure 15-44] Same as above [Figure 15-45] Same as above [Figure 15-46] Same as above [Figure 15-47] Same as above [Figure 15-48] Same as above [Figure 15-49] Same as above [Figure 15-50] Same as above [Figure 15-51] Same as above [Figure 15-52] Same as above [Figure 15-53] Same as above [Figure 15-54] Same as above [Figure 15-55] Same as above [Figure 15-56] Same as above [Figure 15-57] Same as above [Figure 15-58] Same as above [Figure 15-59] Same as above [Figure 15-60] Same as above [Figure 15-61] Same as above [Figure 15-62] Same as above [Figure 15-63] Same as above [Figure 15-64] Same as above [Figure 15-65] Same as above [Figure 15-66] Same as above [Figure 15-67] Same as above [Figure 15-68] Same as above [Figure 15-69] Same as above [Figure 15-70] Same as above [Figure 15-71] Same as above [Figure 15-72] Same as above [Figure 15-73] Same as above [Figure 15-74] Same as above [Figure 15-75] Same as above [Figure 15-76] Same as above [Figure 15-77] Same as above [Figure 15-78] Same as above [Figure 15-79] Same as above [Figure 15-80] Same as above [Figure 15-81] Same as above [Figure 15-82] Same as above [Figure 15-83] Same as above [Figure 15-84] Same as above [Figure 15-85] Same as above [Figure 15-86] Same as above [Figure 15-87] Same as above [Figure 15-88] Same as above [Figure 15-89] Same as above [Figure 15-90] Same as above [Figure 15-91] Same as above [Figure 15-92] Same as above [Figure 15-93] Same as above [Figure 15-94] Same as above [Figure 15-95] Same as above [Figure 15-96] Same as above [Figure 15-97] Same as above [Figure 15-98] Same as above [Figure 15-99] Same as above [Figure 15-100] Same as above [Figure 15-101] Same as above [Figure 15-102] Same as above [Figure 15-103] Same as above [Figure 15-104] Same as above [Figure 15-105] Same as above [Figure 15-106] Same as above [Figure 15-107] Same as above [Figure 15-108] Same as above [Figure 15-109] Same as above [Figure 15-110] Same as above [Figure 15-111] Same as above [Figure 15-112] Same as above [Figure 15-113] Same as above [Figure 15-114] Same as above [Figure 15-115] Same as above [Figure 15-116] Same as above [Figure 15-117] Same as above [Figure 15-118] Same as above [Figure 15-119] Same as above [Figure 15-120] Same as above [Figure 15-121] Same as above [Figure 15-122] Same as above [Figure 15-123] Same as above [Figure 15-124] Same as above [Figure 15-125] Same as above [Figure 15-126] Same as above [Figure 15-127] Same as above [Figure 15-128] Same as above [Figure 15-129] Same as above [Figure 15-130] Same as above

[0053] [Figure 16-1]Figure 16 provides a sequence listing for the transposase proteins described herein (SEQ ID NOS: 10176-10523). The proteins are also represented as amino acid sequences in typical FASTA format, with additional lines for secondary structure predictions for each residue. Additional formatting is used to indicate subsequences of interest. When available, annotations include: dark gray highlights to identify the boundaries of the RuvC-like domain predicted using the DEDD_Tnp_IS110 Pfam domain; light gray highlights to identify the boundaries of the Tnp domain predicted using the Transposase_20 Pfam domain. Bold indicates highly conserved amino acids, with up to five such amino acids in each domain. Secondary structure predictions were generated using the standard mkdssp tool against all available AlphaFold structures of IS110 transposases. These secondary structures were then projected to the sequences in our collection by primary sequence alignment. Different letters indicate the following: H, alphahelix; B, betabridge; E, strand; G, helix_3; I, helix_5; P, polyproline type II helix (Helix_PPII); T, turn; S, bend; and "-", loop. These secondary structures can be used to guide those skilled in the art and identify coiled-coil linker domains. Additional transposase protein sequences are provided in the attached sequence listing as SEQ ID NOS: 10524-20350 and 40357-516430, which are incorporated herein by reference in their entireties. The sequence listing features the start and stop positions of the RuvC-like and Tnp domains, as well as the P1-P5 positions of each domain used in the AlphaFold analysis. [Figure 16-2] Same as above [Figure 16-3] Same as above [Figure 16-4] Same as above [Figure 16-5] Same as above [Figure 16-6] Same as above [Figure 16-7] Same as above [Figure 16-8] Same as above [Figure 16-9] Same as above [Figure 16-10] Same as above [Figure 16-11] Same as above [Figure 16-12] Same as above [Figure 16-13] Same as above [Figure 16-14] Same as above [Figure 16-15] Same as above [Figure 16-16] Same as above [Figure 16-17] Same as above [Figure 16-18] Same as above [Figure 16-19] Same as above [Figure 16-20] Same as above [Figure 16-21] Same as above [Figure 16-22] Same as above [Figure 16-23] Same as above [Figure 16-24] Same as above [Figure 16-25] Same as above [Figure 16-26] Same as above [Figure 16-27] Same as above [Figure 16-28] Same as above [Figure 16-29] Same as above [Figure 16-30] Same as above [Figure 16-31] Same as above [Figure 16-32] Same as above [Figure 16-33] Same as above [Figure 16-34] Same as above [Figure 16-35] Same as above [Figure 16-36] Same as above [Figure 16-37] Same as above [Figure 16-38] Same as above [Figure 16-39] Same as above [Figure 16-40] Same as above [Figure 16-41] Same as above [Figure 16-42] Same as above [Figure 16-43] Same as above [Figure 16-44] Same as above [Figure 16-45] Same as above [Figure 16-46] Same as above [Figure 16-47] Same as above [Figure 16-48] Same as above [Figure 16-49] Same as above [Figure 16-50] Same as above [Figure 16-51] Same as above [Figure 16-52] Same as above [Figure 16-53] Same as above [Figure 16-54] Same as above [Figure 16-55] Same as above [Figure 16-56] Same as above [Figure 16-57] Same as above [Figure 16-58] Same as above [Figure 16-59] Same as above [Figure 16-60] Same as above [Figure 16-61] Same as above [Figure 16-62] Same as above [Figure 16-63] Same as above [Figure 16-64] Same as above [Figure 16-65] Same as above [Figure 16-66] Same as above [Figure 16-67] Same as above [Figure 16-68] Same as above [Figure 16-69] Same as above [Figure 16-70] Same as above [Figure 16-71] Same as above [Figure 16-72] Same as above [Figure 16-73] Same as above [Figure 16-74] Same as above [Figure 16-75] Same as above [Figure 16-76] Same as above [Figure 16-77] Same as above [Figure 16-78] Same as above [Figure 16-79] Same as above [Figure 16-80] Same as above [Figure 16-81] Same as above [Figure 16-82] Same as above

[0054] [Figure 17-1] Figure 17 provides a sequence listing for the donors (SEQ ID NOS: 30354-30529). The donors are represented as 50-nt 5'-3' nucleotide sequences in typical FASTA format with additional formatting to indicate the subsequence of interest. When available, annotations include: light gray highlighting indicating the right end (RE) and left end (LE). In this case, the RE is 5' to the core sequence, and the LE is 3' to the core sequence. The core sequence is represented as single-underlined, unhighlighted text. If present, the programmable portions of the donor corresponding to the LDG and RDG of the bridge RNA are shown in bold. The programmable portion of the donor RE corresponding to the LDG of the bridge RNA is referred to as the left donor (LD), and the programmable portion of the donor LE corresponding to the RDG of the bridge RNA is referred to as the right donor (RD). Additional donor sequences are provided as SEQ ID NOS: 30530-40356 in the attached sequence listing, which are incorporated herein by reference in their entirety. The sequence listing includes the start and stop positions of the core, LE, and RE sequences as features of this sequence listing. [Figure 17-2] Same as above [Figure 17-3] Same as above [Figure 17-4] Same as above [Figure 17-5] Same as above [Figure 17-6] Same as above [Figure 17-7] Same as above [Figure 17-8] Same as above [Figure 17-9] Same as above

[0055] [Figure 18-1]Figure 18 provides a sequence listing of the targets (SEQ ID NOS: 20351-20526). The targets are represented as 50-nt 5'-3' nucleotide sequences in typical FASTA format with additional formatting to indicate the subsequence of interest. When available, annotations include: light gray highlighting indicating the left flank (LF) and right flank (RF). In this case, the LF is 5' to the core sequence, and the RF is 3' to the core sequence. The core sequence is represented as single-underlined, unhighlighted text. The programmable portions of the targets corresponding to the LTG and RTG of the bridge RNA are shown in bold. The programmable portion of the target LF corresponding to the LTG of the bridge RNA is referred to as the left target (LT), and the programmable portion of the donor RF corresponding to the RTG of the bridge RNA is referred to as the right target (RT). Additional target sequences are provided as SEQ ID NOS: 20527-30353 in the attached sequence listing, which are incorporated herein by reference in their entirety. The sequence listing includes the start and stop positions of the core, LF, and RF sequences as features of this sequence listing. [Figure 18-2] Same as above [Figure 18-3] Same as above [Figure 18-4] Same as above [Figure 18-5] Same as above [Figure 18-6] Same as above [Figure 18-7] Same as above [Figure 18-8] Same as above [Figure 18-9] Same as above

[0056] [Figure 19-1]Figure 19 provides the consensus sequence and structure of the bridge RNA sequence (SEQ ID NOs: 795156, 795304, 795287, 795305 to 795328, 795297, 795329, 795330 to 795344, 795289, 795345 to 795351, 795294, 795352 to 795400, 795291, 795401 to 795412, 795300, 795288, 795302, 795413 to 795427, 795301, 795428 to 795440, 7952 96, 795441–795446, 795295, 795447, 795290, 795448–795454, 795298, 795455–795459, 795292, 795460–795468, 795293, 795469, 795470–795471, 795370, 795472–795508, 795299, 795509, 795510–795514, 795303, 795515–795542, 795536, and 795543–795564, respectively, in order of appearance. The name of each model is specified by a line beginning with ">", as in a typical FASTA file. The next line is the consensus sequence for the model, where "n" represents any nucleotide, "R" represents an A or G nucleotide, "Y" represents a C or U nucleotide, and A, C, G, and T represent individual nucleotides. The next four lines show four possible RNA secondary structures using different confidence thresholds when running the ConsAliFold RNA structure prediction algorithm. These four lines correspond to gamma parameters of 4, 8, 16, and 32, respectively, with increasing gamma values ​​representing more permissive models (allowing more structures). The notation used for secondary structures is called "dot-bracket" notation, where matching brackets "(" and ")" indicate base pairs and unpaired bases are indicated by a dot ("."). [Figure 19-2] Same as above [Figure 19-3] Same as above [Figure 19-4] Same as above [Figure 19-5] Same as above [Figure 19-6] Same as above [Figure 19-7] Same as above [Figure 19-8]Same as above [Figure 19-9] Same as above [Figure 19-10] Same as above [Figure 19-11] Same as above [Figure 19-12] Same as above [Figure 19-13] Same as above [Figure 19-14] Same as above [Figure 19-15] Same as above [Figure 19-16] Same as above [Figure 19-17] Same as above [Figure 19-18] Same as above [Figure 19-19] Same as above [Figure 19-20] Same as above [Figure 19-21] Same as above [Figure 19-22] Same as above [Figure 19-23] Same as above [Figure 19-24] Same as above [Figure 19-25] Same as above [Figure 19-26] Same as above [Figure 19-27] Same as above [Figure 19-28] Same as above [Figure 19-29] Same as above [Figure 19-30] Same as above [Figure 19-31] Same as above [Figure 19-32] Same as above [Figure 19-33] Same as above [Figure 19-34] Same as above [Figure 19-35] Same as above [Figure 19-36] Same as above [Figure 19-37] Same as above [Figure 19-38] Same as above [Figure 19-39] Same as above [Figure 19-40] Same as above [Figure 19-41] Same as above [Figure 19-42] Same as above [Figure 19-43] Same as above [Figure 19-44] Same as above [Figure 19-45] Same as above [Figure 19-46] Same as above [Figure 19-47] Same as above [Figure 19-48] Same as above [Figure 19-49] Same as above [Figure 19-50] Same as above [Figure 19-51] Same as above [Figure 19-52] Same as above [Figure 19-53] Same as above [Figure 19-54] Same as above [Figure 19-55] Same as above [Figure 19-56] Same as above [Figure 19-57] Same as above [Figure 19-58] Same as above [Figure 19-59] Same as above [Figure 19-60] Same as above [Figure 19-61] Same as above [Figure 19-62] Same as above [Figure 19-63] Same as above

[0057] [Figure 20-1] Figure 20 shows the AlphaFold model of IS621 transposase used in the structural analysis. All available AlphaFold structures of IS110 transposases were aligned to this model using the TM-align algorithm to generate a TM score. This analysis established that a TM score cutoff of 0.5 is both sensitive and accurate for identifying IS110 transposases. [Figure 20-2] Same as above [Figure 20-3] Same as above [Figure 20-4] Same as above [Figure 20-5] Same as above [Figure 20-6] Same as above [Figure 20-7] Same as above [Figure 20-8] Same as above [Figure 20-9] Same as above [Figure 20-10] Same as above [Figure 20-11] Same as above [Figure 20-12] Same as above [Figure 20-13] Same as above [Figure 20-14] Same as above [Figure 20-15] Same as above [Figure 20-16] Same as above [Figure 20-17] Same as above [Figure 20-18] Same as above [Figure 20-19] Same as above [Figure 20-20] Same as above [Figure 20-21] Same as above [Figure 20-22] Same as above [Figure 20-23] Same as above [Figure 20-24] Same as above [Figure 20-25] Same as above [Figure 20-26] Same as above [Figure 20-27] Same as above [Figure 20-28] Same as above [Figure 20-29] Same as above [Figure 20-30] Same as above [Figure 20-31] Same as above [Figure 20-32] Same as above [Figure 20-33] Same as above [Figure 20-34] Same as above [Figure 20-35] Same as above [Figure 20-36] Same as above [Figure 20-37] Same as above [Figure 20-38] Same as above [Figure 20-39] Same as above [Figure 20-40] Same as above [Figure 20-41] Same as above [Figure 20-42] Same as above [Figure 20-43] Same as above [Figure 20-44] Same as above [Figure 20-45] Same as above [Figure 20-46] Same as above [Figure 20-47] Same as above [Figure 20-48] Same as above [Figure 20-49] Same as above [Figure 20-50] Same as above [Figure 20-51] Same as above [Figure 20-52] Same as above [Figure 20-53] Same as above [Figure 20-54] Same as above [Figure 20-55] Same as above [Figure 20-56] Same as above [Figure 20-57] Same as above [Figure 20-58] Same as above [Figure 20-59] Same as above [Figure 20-60] Same as above [Figure 20-61] Same as above [Figure 20-62] Same as above [Figure 20-63] Same as above [Figure 20-64] Same as above [Figure 20-65] Same as above [Figure 20-66] Same as above [Figure 20-67] Same as above [Figure 20-68] Same as above [Figure 20-69] Same as above [Figure 20-70] Same as above [Figure 20-71] Same as above [Figure 20-72] Same as above [Figure 20-73] Same as above [Figure 20-74] Same as above [Figure 20-75] Same as above [Figure 20-76] Same as above [Figure 20-77] Same as above [Figure 20-78] Same as above [Figure 20-79] Same as above [Figure 20-80] Same as above [Figure 20-81] Same as above [Figure 20-82] Same as above [Figure 20-83] Same as above [Figure 20-84] Same as above [Figure 20-85] Same as above [Figure 20-86] Same as above [Figure 20-87] Same as above [Figure 20-88] Same as above [Figure 20-89] Same as above [Figure 20-90] Same as above [Figure 20-91] Same as above [Figure 20-92] Same as above [Figure 20-93] Same as above [Figure 20-94] Same as above [Figure 20-95] Same as above [Figure 20-96] Same as above [Figure 20-97] Same as above [Figure 20-98] Same as above [Figure 20-99] Same as above

[0058] [Figure 21-1] FIG. 21 shows the RuvC-like DEDD catalytic domain motif of IS110 transposases belonging to the IS110 group. [Figure 21-2] Same as above [Figure 21-3] Same as above [Figure 21-4] Same as above [Figure 21-5] Same as above [Figure 21-6] Same as above

[0059] [Figure 22] Figure 22 shows the motif of the "D" region of the canonical DEDD catalytic motif of IS110 transposases belonging to the IS110 group. SEQ ID NOs are shown in parentheses.

[0060] [Figure 23-1] Figure 23 shows the motif of the "E" region of the canonical DEDD catalytic motif of IS110 transposases belonging to the IS110 group. SEQ ID NOs are shown in parentheses. [Figure 23-2] Same as above

[0061] [Figure 24-1] Figure 24 shows the motif of the "DD" region of the canonical DEDD catalytic motif of IS110 transposases belonging to the IS110 group. SEQ ID NOs are shown in parentheses. [Figure 24-2] Same as above

[0062] [Figure 25-1] Figure 25 shows the RuvC-like DEDD catalytic domain motif of IS110 transposase, which belongs to the IS1111 group. [Figure 25-2] Same as above [Figure 25-3] Same as above [Figure 25-4] Same as above

[0063] [Figure 26] Figure 26 shows the motif of the "D" region of the canonical DEDD catalytic motif of IS110 transposases belonging to the IS1111 group. SEQ ID NOs are shown in parentheses.

[0064] [Figure 27-1] Figure 27 shows the motif of the "E" region of the canonical DEDD catalytic motif of IS110 transposases belonging to the IS1111 group. SEQ ID NOs are shown in parentheses. [Figure 27-2] Same as above [Figure 27-3] Same as above

[0065] [Figure 28-1] Figure 28 shows the motif of the "DD" region of the canonical DEDD catalytic motif of IS110 transposases belonging to the IS1111 group. SEQ ID NOs are shown in parentheses. [Figure 28-2] Same as above

[0066] [Figure 29-1] FIG. 29 shows the transposase domain motifs of IS110 transposases belonging to the IS110 group. [Figure 29-2] Same as above [Figure 29-3] Same as above [Figure 29-4] Same as above [Figure 29-5] Same as above [Figure 29-6] Same as above [Figure 29-7] Same as above [Figure 29-8] Same as above [Figure 29-9] Same as above

[0067] [Figure 30-1] Figure 30 shows the motif of the first conserved region of the transposase domain of IS110 transposases belonging to the IS110 group. SEQ ID NOs are shown in parentheses. [Figure 30-2] Same as above

[0068] [Figure 31-1] Figure 31 shows the motif of the second conserved region of the transposase domain of IS110 transposases belonging to the IS110 group. SEQ ID NOs are shown in parentheses. [Figure 31-2] Same as above

[0069] [Figure 32-1] Figure 32 shows the transposase domain motif of IS110 transposase, which belongs to the IS1111 group. [Figure 32-2] Same as above [Figure 32-3] Same as above [Figure 32-4] Same as above [Figure 32-5] Same as above

[0070] [Figure 33-1]Figure 33 shows the motif of the first conserved region of the transposase domain of IS110 transposases belonging to the IS1111 group. SEQ ID NOs are shown in parentheses. [Figure 33-2] Same as above

[0071] [Figure 34-1] Figure 34 shows the motif of the second conserved region of the transposase domain of IS110 transposases belonging to the IS1111 group. SEQ ID NOs are shown in parentheses. [Figure 34-2] In Figures 21-34, motifs are in the general Prosite format, where x represents any amino acid, x(n) represents any n amino acids, and x(n,m) represents any n to m amino acids. The order is determined by the frequency of occurrence of the domain in the transposase sequence database. In Figures 22-24, 26-28, 30-31, and 33-34, motifs are shown as a semicolon-separated list of each motif.

[0072] [Figure 35A]Figures 35A-B show additional examples of predicted bridge RNA secondary structures with predicted LTG, RTG, LDG, and RDG guide sequences. (A) Schematic diagrams of six bridge RNA consensus structures derived from three IS110 group elements and three IS1111 group elements. IS110 group elements typically encode a bridge RNA at the 5' non-coding end (LE) of the element, whereas IS1111 group elements typically encode a bridge RNA at the 3' non-coding end (RE). Guide sequences are color-coded according to the sequence to which they bind: target (blue), donor (orange), or core (green). In some members of the IS1111 group, donor-binding guide sequences are often found within large multi-loop structures rather than internal loops. (B) A more detailed representation of the same structure and sequence as in (A). The consensus secondary structure is shown with IUPAC nucleotide code circles color-coded according to conservation. Highlighted guide sequences are displayed above the corresponding target (SEQ ID NOS: 798527, 798529, 798531, 798533, 798535, 798537, in order of appearance) and donor (SEQ ID NOS: 798528, 798530, 798532, 798534, 798536, 798538, in order of appearance) sequences for comparison. LTG, RTG, LDG, and RDG are directly labeled. The bridge RNA structure is a representation of SEQ ID NOS: 795344, 795370, 795303, 795295, 795293, and 795287, excluding gap positions and trimming excess unstructured bases. [Figure 35B-1] Same as above [Figure 35B-2] Same as above [Figure 35B-3] Same as above [Figure 35B-4] Same as above [Figure 35B-5] Same as above [Figure 35B-6] Same as above

[0073] [Figure 36A]Figures 36A-E demonstrate the utility of extending the natural length of the right targeting guide (RTG) to enhance the efficiency and specificity of programmable recombination. (A) Schematic illustrating how longer RTGs can be reprogrammed, as well as how cores can be reprogrammed in combination with longer RTG reprogramming. (B) Relative recombination rates between reprogrammed cores and donors with 4- or 7-bp homologous RTGs and targets. The assay detailed in Figure 5D was used. The results show that having longer RTG homology enhances the efficiency of recombination with WT and reprogrammed core sequences. (C) Schematic of the approach for genome integration, identical in approach to Figure 9A. (D) On- and off-target integration frequencies using 4- or 7-base RTGs for targeting. The same bridge RNA as in Figure 9B was utilized to integrate donor cargo into the E. coli genome using either 4- or 7-base RTGs. Integration sites were binned by the number of differences from the 11-bp target site sequence intended by the bridge RNA programmed with a 4-nt RTG. (E) Rank order of integration sites averaged over two replicates. The same data as in (D) is illustrated by relative integration counts. High-frequency integrations are highlighted by graphically depicting their sequences, demonstrating how relative targeting specificity is altered when comparing 4-nt and 7-nt RTGs. [Figure 36B] Same as above [Figure 36C] Same as above [Figure 36D] Same as above [Figure 36E] Same as above

[0074] [Figure 37A]Figures 37A-E show evaluation of donor boundaries for the IS110 bridge recombinase system. (A) Schematic for assaying sequence preference upstream of the donor sequence. Six nucleotides upstream of the LD are altered. Recombinants are selected using kanamycin resistance, and successful recombinants are determined by next-generation sequencing (NGS). The target is SEQ ID NO: 798549, and the donor is SEQ ID NO: 798547. (B) Schematic for assaying sequence preference downstream of the donor sequence. Eight nucleotides downstream of the fourth position of the RD, including part of the RD, are altered. Assay parameters are otherwise identical to those shown in A. The target is SEQ ID NO: 798549, and the donor is SEQ ID NO: 798548. (C) Nucleotide requirements upstream and downstream of the donor sequence. The 5' and 3' STIR sequences are highlighted in pink. This figure shows SEQ ID NO: 795196. (D-E) Sequence preference upstream (D) and downstream (E) of the donor sequence. The 5' and 3' STIR sequences are highlighted in pink. [Figure 37B] Same as above [Figure 37C] Same as above [Figure 37D] Same as above [Figure 37E] Same as above

[0075] [Figure 38A]Figures 38A-C show plasmid-plasmid recombination in human cells. (A) Schematic of the plasmid-plasmid recombination assay in human cells. pEffector expresses the bridge RNA and recombinase from the U6 and Ef1a promoters, respectively. pDonor and pTarget are recombined by cotransfection with pEffector. Recombination is detected by PCR of the LT-RD junction using primers F and R. (B) Verification of plasmid-plasmid recombination. PCR of the LT-RD junction is performed using pDonor and pTarget alone, and pDonor, pTarget, and pEffector, with pEffector lacking the bridge RNA. The recombinase on pEffector was evaluated with three different NLS formats. The recombinase shown is IS621 recombinase with a bridge RNA specific for its wild-type donor sequence and the reprogrammed target sequence, Target 01. The target-binding loop RTG encodes 7 bp of homology to the target. (C) Sanger sequencing confirmation of recombination. Sanger sequencing traces are aligned to the entire PCR at the LT-RD junction (top), and a zoomed-in view shows the nucleotides near the LT-RD (bottom). The figure shows SEQ ID NOs: 798550, 798550, and 798551 in order of appearance. [Figure 38B] Same as above [Figure 38C-1] Same as above [Figure 38C-2] Same as above

[0076] [Figure 39A]Figure 39A-D shows plasmid inversion in human cells using various orthologs. (A) Schematic of the plasmid inverse recombination assay in human cells. pEffector expresses a bridge RNA and a recombinase from the U6 and Ef1a promoters, respectively. The recombinase is fused to the P2A self-cleaving peptide and EGFP, and recombinase expression is measured. Recombination is detected by PCR of the LT-RD junction using primers F and R. (B) Verification of plasmid inverse recombination. PCR of the LT-RD junction is performed with three different NLS configurations for IS621 23122 recombinase in the presence and absence of bridge RNA. (C) Percentage of cells expressing EGFP 72 hours after transfection. Four IS110 orthologs with different NLS configurations are shown. (D) Percentage of mCherry-positive cells within the EGFP-positive cell population. Four IS110 orthologs with different NLS configurations are shown. For each ortholog, the WT target and donor sequences were recombined. For IS621_127209 and IS621_23122, the sequences flanking the WT 4-nt RT were modified to allow 7 bp between the RT and the WT RTG. [Figure 39B] Same as above [Figure 39C] Same as above [Figure 39D] Same as above

[0077] [Figure 40A]Figures 40A-D show bridge RNA engineering to improve efficiency and specificity. (A) Schematic of IS110 element IS621_23122 showing the approximate boundary locations of the bridge RNA. A 179-nt bridge RNA (bRNA179) spans from the start of the bridge RNA to the end of the LE of the element. A 260-nt bridge RNA (bRNA260) starts at the same location and extends into the CDS of the recombinase. (B) Bridge editing efficiency of an inverted reporter using bridge RNAs of different lengths. Extending the bridge RNA to its natural sequence context of 260 nt increases efficiency compared to the 179-nt bridge RNA. (C) Schematic comparing the WT target binding loop with the LTG-shifted target binding loop. The LTG shift improves specificity by allowing targeting of 16-nt target sequences by binding to the 9 bp preceding the core rather than the 9 bp containing the core. The target is SEQ ID NO: 798552, and the donor is SEQ ID NO: 798553. (D) Bridge editing efficiency of the inverted reporter using WT bridge RNA and LTG-shifted bridge RNA, both of which utilize the additional 81 nt added to the 3' end of the bridge RNA in panel b. [Figure 40B] Same as above [Figure 40C] Same as above [Figure 40D] Same as above

[0078] [Figure 41A]Figures 41A-C show engineering of the human genome by delivery of a large DNA cargo and a bridge editor. (A) Schematic of bridge editing of the human genome by delivery of a donor plasmid. A recombinase and bridge RNA specific for the donor plasmid (pDonor, 4.8 kb) and a target sequence within the genome result in integration of the donor into the genome. Recombination is detected by PCR of the LT-RD junction using primers F and R. (B) PCR detection of the LT-RD junction from genomic DNA. (C) Sanger sequencing confirmation of donor integration into the human genome. Sanger sequencing traces are aligned to the entire PCR of the LT-RD junction (top), and a zoomed-in view shows the nucleotides adjacent to the LT-RD (bottom). The figure shows SEQ ID NOs: 798554, 798555, and 798556 in order of appearance. [Figure 41B] Same as above [Figure 41C-1] Same as above [Figure 41C-2] Same as above

[0079] [Figure 42A]Figures 42A-D show engineering of the human genome by delivery of a bridge editor alone. (A) Schematic of bridge editing of the human genome for an inversion by delivery of only a recombinase and bridge RNA. Recombinase and bridge RNA specific to the genome donor and genome target sequences result in an inversion when the donor and target are on opposite strands. Recombination is detected by PCR of the RD-LT junction using primers L and L' and PCR of the LD-RT junction using primers R and R'. Various orientations of the target and donor result in an inversion, one of which is shown here. (B-C) PCR detection of RD-LT and LD-RT of four different bridge RNAs on an agarose gel. The chromosomal locus targeted by the bridge RNA is shown (top), and the relative orientations of the donor and target before and after recombination are shown (bottom). (D) Example of Sanger sequencing confirmation of the inversion locus from panel b. Sanger sequencing traces are aligned to the entire PCR at the RD-LT junction (top left) and zoomed in to show the nucleotides adjacent to RD-LT (bottom left). Sanger sequencing traces are aligned to the entire PCR at the LD-RT junction (top right) and zoomed in to show the nucleotides adjacent to LD-RT (bottom right). The figure shows SEQ ID NOs: 798557, 798557, 798558, 798559, and 798558 in order of appearance. (E) Schematic of bridge editing of the human genome for excision by delivery of recombinase and bridge RNA alone. Recombinase and bridge RNA specific for the genome donor and genome target sequences result in excision when the donor and target are on the same strand. PCR of the LD-RT junction with primers G and G' detects excision from the locus, while PCR of the LT-RD junction with primers E and E' detects excised DNA. Various orientations of the target and donor result in inversions, one of which is shown here. [Figure 42B] Same as above [Figure 42C] Same as above [Figure 42D-1] Same as above [Figure 42D-2] Same as above [Figure 42D-3] Same as above [Figure 42D-4] Same as above [Figure 42E] Same as above

[0080] [Figure 43A]Figures 43A-F show the operation of the split-bridge RNA system for recombination. (A) Schematic of a recombination assay using donor- and target-specific LE-encoded bridge RNAs. The donor is SEQ ID NO: 795177. (B) Schematic of a recombination assay using an LE-encoded bridge RNA and separately expressed target-binding loops (TBLs). The target-binding loops of the LE-encoded bridge RNA were inactivated by reprogramming LTG and RTG to have no complementarity to any sequences in the plasmid or organism, while the donor-binding loop (DBL) is specific for the donor site sequence. (C) Comparison of recombination efficiency using the WT bridge RNA structure (A) and the split-bridge RNA depicted in (B). (D) Schematic of various split-bridge RNA systems. The bridge RNA is depicted as a sequence in which one or the other binding loop has been reprogrammed to have no specificity for sequences found in the system. Some versions completely split the bridge RNA into two halves, with separate TBLs and DBLs. Some versions eliminate additional accessory or unstructured nucleotides. The right panel shows the approximate specificity and structure of various iterations of the split-bridge RNA system. (E) Insertion of cargo into the lacZ gene of the E. coli genome. Insertion is achieved using a bridge RNA with a TBL targeting the lacZ gene. Blue-white screening for lacZ activity using beta-galactosidase is used to select colonies with genomic insertions. PCR and agarose gel (right) confirm integration of the cargo into the genome. (F) Insertion of cargo into the lacZ gene of the E. coli genome. Insertion is achieved using a plasmid system or a bridge RNA with a TBL reprogrammed to have no specificity for sequences within E. coli, while a separate TBL specific for the lacZ gene is expressed from a synthetic promoter, such as the system illustrated in (B). Blue-white screening for lacZ activity using beta-galactosidase is used to select colonies with genomic insertions. PCR and agarose gel (right) confirm integration of the cargo into the genome. [Figure 43B] Same as above [Figure 43C] Same as above [Figure 43D] Same as above [Figure 43E] Same as above [Figure 43F] Same as above

[0081] [Figure 44A]Figures 44A-B show a summary of the mismatch tolerance between the IS110 bridge RNA target binding loop and its target. (A) Schematic of the antibiotic resistance reporter design. A minimal donor (22 bp) is encoded on a plasmid adjacent to a kanamycin resistance gene. A second plasmid encodes the target, bridge RNA, and transposase. The target is linked to the bridge RNA using a barcode. Recombination between the donor and target plasmids results in survival of the E. coli, and functional bridge RNA target loop-target pairs are recorded using next-generation sequencing (left). Schematic of the target specificity screen design. The target and target loop are varied, except for the target core and the LTG and RTG subsequences that bind to the core. The donor loop (and donor) are kept constant. Target and target loop pairs are designed to assay single mismatches, double mismatches, and total mismatches. Targets in the screen are selected to reduce the number of off-targets in the E. coli genome. (B) Sequence abundance of target and target loop pairs. Abundance is measured by barcode counts per million reads. Target / target loop pairs with 0 mismatches are generally abundant, and increasing the number of mismatches decreases abundance. (D) Sequence logos of top quintile targets. The relative enrichment of nucleotides at each target position is shown for the top quintile of 6,364 target / target loop pairs with 0 mismatches. (B) Mismatch tolerance at each position of the 11-bp target sequence. The x-axis indicates the target position, holding the CT core constant. The top panel shows the target nucleotide recovery frequency as a percentage of recombinants recovered at each position when the target binding loop contains an A at each guide position. The second, third, and fourth panels show the same for when the target binding loop contains a C, G, or U at each position. Target positions are shown as the top strand of DNA. [Figure 44B] Same as above DETAILED DESCRIPTION OF THE INVENTION

[0082] The present invention relates to the IS110 transposon family. IS110 transposons encode both a "bridge RNA" molecule and a transposase protein. The bridge RNA molecule, in cooperation with the transposase, mediates site-specific recombination between one or more DNA molecules containing a target site sequence and a donor site sequence. The target site sequence and the donor site sequence may be on the same DNA molecule or on different DNA molecules. Generally, as used herein, the target site and the donor site sequence simply refer to nucleic acid sequences that associate with or are recognized by the IS110 bridge RNA and the transposase complex. Depending on the orientation of these sequences and whether they are on the same molecule or different molecules, a transposition reaction can result in either insertion (or translocation), excision recombination, or inversion between the target site sequence and the donor site sequence. Thus, in the case of an insertion or translocation reaction, the target sequence and the donor site sequence are on different molecules. In excision recombination and inversion, the target site sequence and donor site sequence are on the same molecule, and depending on the orientation of the target site sequence and donor site sequence, the intervening sequence is excised or inverted. Such recombination reactions can be used to programmably recombine any DNA sequence with any other DNA sequence, without the need to use DNA sequences derived from IS110 elements. More specifically, the present invention provides recombinant IS110 transposons whose encoded bridge RNA molecules are programmable by modifying sequences within the target and / or donor binding loops of the bridge RNA, thereby engineering the bridge RNA to specifically bind to a sequence of interest. In one category of programmable transposition, the bridge RNA is designed to recombine a desired donor DNA molecule with a desired target DNA molecule to insert a sequence located on a different DNA molecule or translocate a sequence on a different DNA molecule. In another category of programmable transposition, the bridge RNA is designed to recombine a desired donor DNA sequence with a desired target DNA molecule to excise or invert an intervening sequence located on the same DNA molecule.Additionally, the present invention encompasses non-programmed uses of IS110 family transposons. For example, a non-programmed IS110 bridge RNA (in which the target and donor binding loops have not been modified to alter the binding specificity of the bridge RNA) and transposase complex can be used to target naturally occurring target and donor site sequences in prokaryotic genomes, naturally occurring target and donor site sequences in eukaryotic genomes, target and donor site sequences introduced into prokaryotic genomes, and target and donor site sequences introduced into eukaryotic genomes.

[0083] A. Definition The terms "polynucleotide," "nucleotide sequence," "nucleic acid," "nucleic acid molecule," "nucleic acid segment," and "oligonucleotide" are used interchangeably. They refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof. A polynucleotide can be any three-dimensional structure. Polynucleotides may have any structure and may perform any function, known or unknown. The following are non-limiting examples of polynucleotides: coding or non-coding regions of a gene or gene fragment, loci defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, small interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Polynucleotides may contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. The sequence of nucleotides may be interrupted by non-nucleotide components. Polynucleotides may be further modified after polymerization, such as by conjugation with a labeling component.

[0084] The term "DNA" refers to deoxyribonucleic acid, including, but not limited to, genomic or non-genomic DNA present in a cell, or isolated forms of such DNA. Genomic or non-genomic DNA includes, but is not limited to, chromosomal or non-chromosomal DNA, such as episomal DNA, viral DNA, plasmid DNA, mitochondrial DNA, cellular DNA, or chloroplast DNA.

[0085] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymers of amino acids of any length. The polymers may be linear or branched, may comprise modified amino acids, and may be interrupted by non-amino acids. These terms also encompass amino acid polymers that have been modified, for example, by disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or other manipulations, such as conjugation with a labeling component. As used herein, the term "amino acid" includes natural and / or unnatural or synthetic amino acids, including glycine and both the D- and L-optical isomers, as well as amino acid analogs and peptidomimetics.

[0086] As used herein, the term "nuclease" refers to an agent, such as a protein or small molecule, that can cleave the phosphodiester bond that links nucleotide residues in a nucleic acid molecule. In some embodiments, a nuclease is an enzyme that can bind to a nucleic acid molecule and cleave the phosphodiester bond that links nucleotide residues within a nucleic acid molecule. A nuclease can be an endonuclease that cleaves phosphodiester bonds within a polynucleotide chain, or an exonuclease that cleaves phosphodiester bonds at the end of a polynucleotide chain. In some embodiments, a nuclease is a site-specific nuclease that binds to and / or cleaves specific phosphodiester bonds within a specific nucleotide sequence. The term "nickase" refers to an endonuclease that cleaves only one strand of a DNA duplex.

[0087] The term "excisionase" as used herein refers to a host-derived, bacteriophage, or mobile genetic element sequence-specific DNA-binding protein that is involved in removing DNA from a nucleotide sequence and repairing the DNA with or without traces of the sequence. The removed DNA can be in the form of linear or circular ssDNA or dsDNA.

[0088] "Sequence-specific" refers to, but is not limited to, recombination or recombination events that occur at predictable loci or identifiable nucleotide sequences, or modifications of nucleotides at predetermined sequence locations.

[0089] As used herein, the term "transposon" refers to a polynucleotide (or nucleic acid segment) that can be copied or moved into a new nucleic acid sequence context by the action of a transposase. An insertion sequence (IS) element refers to a transposon that encodes the minimal components required for recombination of nucleotide sequences (e.g., a transposase and a bridge RNA). IS elements may be referred to herein as IS elements, IS110 elements, or transposons.

[0090] As used herein, the term "transposase" refers to an enzyme that is a component of a functional nucleic acid-protein complex (e.g., a transpososome) capable of transposition and mediates transposition. A transposase can be composed of a single protein or multiple proteins. A transposase can be an enzyme that can form a functional complex with a transposon end, a transposon end sequence, or a transposon-derived sequence. In certain embodiments, the term "transposase" can refer to an integrase, a recombinase, an invertase, or an excisionase. Described herein are transposases derived from the IS110 family of transposons. The IS110 "transposases" described herein are also referred to as IS110 "recombinases."

[0091] As used herein, the term "transposition reaction" refers to a reaction, such as an integration reaction, recombination reaction, inversion reaction, or excision reaction, in which a transposase recombines a DNA polynucleotide containing a donor site sequence with a DNA polynucleotide containing a target site sequence. A transposition reaction can occur when the donor site sequence and the target site sequence are on one or more DNA molecules. Both the target site sequence and the donor site sequence can contain sequences or secondary structures. The target site sequence and the donor site sequence can contain sequences or secondary structures recognized by the transposase and / or the insertion motif sequence, in which case the transposase cleaves or makes staggered nicks in the target polynucleotide sequence.

[0092] As used herein, the term "transposon end sequence" refers to a nucleotide sequence at the distal end of a transposon. A transposon end sequence, or a subsequence thereof, can be a DNA sequence that is recognized by a transposase to form a transpososome complex and carry out the transposition reaction. In certain embodiments described herein, the transposon end sequence is derived from the non-coding end sequence of an IS110 family transposon.

[0093] The practice of aspects of the present invention may employ, unless otherwise indicated, conventional techniques of cell biology, cell culture, molecular biology, transgenic biology, microbiology, recombinant DNA, and biochemistry that are within the skill of the art, and such techniques are fully explained in the literature. For example, Sambrook (2001), Fritsch and Maniatis, eds., Molecular Cloning A Laboratory Manual, 3rd Ed., (Cold Spring Harbor Laboratory Press: 1989), DNA Cloning, Volumes I and II (DN Glover ed., 1985), Oligonucleotide Synthesis (MJ Gait ed., 1984), Mullis et al. al., U.S. Patent No. 4,683,195; Nucleic Acid Hybridization (BD Hames & SJ Higgins eds. 1984), Transcription and Translation (BD Hames & SJ Higgins eds. 1984), Culture Of Animal Cells (RI Freshney, Alan R. Liss, Inc., 1987), Immobilized Cells and Enzymes (IRL Press, 1986), B. Perbal, A Practical Guide To Molec ular Cloning (1984), the series, Methods In Enzymology (Academic Press, Inc., NY), specifically, Methods In Enzymology, Vols. 154 and 155 (Wu et al. eds.), Gene Transfer Vectors For Mammalian Cells (JH Miller and MP Calos eds., 1987, Cold Spring Harbor Laboratory), Immunochemical Methods In Cell And Molecular Biology (Caner and Walker, eds., Academic Press, London, 1987), Handbook Of Experimental Immunology, Volumes I-IV (DM Weir and CC Blackwell, eds., 1986), Manipulating the Mouse Embryo, (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1986), and subsequent editions thereof.

[0094] One skilled in the art can obtain proteins in several ways, including, but not limited to, isolating the protein by biochemical means or expressing a nucleotide sequence encoding the protein of interest by genetic engineering methods.

[0095] Proteins are encoded by nucleic acids (including, for example, genomic DNA, messenger RNA (mRNA), complementary DNA (cDNA), synthetic DNA, as well as any form of corresponding RNA). Nucleic acids encoding proteins can be produced by recombinant DNA technology, and such recombinant nucleic acids can be prepared by conventional techniques, including chemical synthesis, genetic engineering, enzymatic technology, or a combination thereof.

[0096] B. IS110 element The IS110 family of transposons refers to a family of transposons widely distributed in prokaryotic genomes. They are classified into two groups, the IS110 group and the IS1111 group, which cumulatively encode transposases that exhibit various insertion site specificities. In addition to their transposase activity, IS110 transposases can also exhibit invertase and excisionase activities.

[0097] The life cycle of an IS110 element is shown in Figure 1C. A linear IS110 element integrated into a target site contains a left non-coding end (LE), a coding sequence for a transposase (Tpase), and a right non-coding end (RE). In some embodiments, the IS110 element is flanked by repeated core sequences, as shown in Figure 1C. The IS110 element excises itself, resulting in a pre-insertion ("target") site bearing an LF-core (if present)-RF and a circular element with an RE-core (if present)-LE-Tpase. The ligation of the RE-LE junction forms a "donor" site sequence as a subsequence of the RE-LE junction, which includes the other core sequence found in the integrated element, if present. The donor site sequence may also contain a subterminal inverted repeat (STIR). The ligation of the RE-LE may form a promoter, which, in the appropriate cellular context, can drive expression from the LE or RE of an RNA molecule, referred to herein as a bridge RNA. The promoter can also drive expression of the transposase in the appropriate cellular context. The bridge RNA encoded within the LE or RE forms an RNA-protein complex with the transposase, recognizing the donor site and / or target site sequence and mediating transposition. The circular form of the element can reinsert into the target site or any other target site sequence recognized by the bridge RNA-transposase complex.

[0098] The left non-coding end (LE) of an IS110 element refers to the nucleotide sequence 5' to the start codon of the IS110 element-encoding IS110 transposase, extending up to (upstream from) the 5' end of the core or element. Thus, the LE is simply defined as the sequence between the CDS and the 5' end of the core or element. The 5' end of an element can be defined using comparative (meta)genomics, by analysis of bridge RNA specificity for donor sequences found at the end of the LE, or by BLAST similarity searches against IS110 sequences defined by the previous two methods. See Examples 1 and 2. In some embodiments, the LE comprises the LE sequence shown in Figure 15 (SEQ ID NOS: 1-348) or Figure 17 (SEQ ID NOS: 30354-30529). In some embodiments, the LE comprises the LE sequence shown in SEQ ID NOS: 349-10175 or 30530-40356.

[0099] The right non-coding end (RE) of an IS110 element refers to the nucleotide sequence 3' to the stop codon of the IS110 element-encoding IS110 transposase, extending downstream to the 3' end of the core or element. Thus, the RE is simply defined as the sequence between the CDS and the 3' end of the core or element. The 3' end of an element can be defined using comparative (meta)genomics, by analysis of bridge RNA specificity for donor sequences found at the ends of the RE, or by BLAST similarity searches against IS110 sequences whose ends are defined by the previous two methods. See Examples 1 and 2. In some embodiments, the RE comprises the RE sequence shown in Figure 15 (SEQ ID NOS: 1-348) or Figure 17 (SEQ ID NOS: 30354-30529). In some embodiments, the RE comprises the RE sequence shown in SEQ ID NOS: 349-10175 or 30530-40356.

[0100] For IS110 transposons containing a core sequence, the core refers to the identical nucleotide sequences found immediately 5' and 3' to the left non-coding end (LE) and right non-coding end (RE), respectively. The core was previously referred to as the "target-intervening core" or "TIC," and any reference to the target-intervening core or TIC refers to the core sequence. In some embodiments, the core sequence is 1-10 nucleotides in length. In some embodiments, the core sequence is 1-5 nucleotides in length. In some embodiments, the core sequence is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides in length. In particular embodiments, the core sequence is 2 nucleotides in length. In some embodiments, the core comprises a core sequence shown in Figure 15 (SEQ ID NOS: 1-348) or Figure 17 (SEQ ID NOS: 30354-30529). In some embodiments, the core comprises a core sequence set forth in SEQ ID NOs: 349-10175 or 30530-40356.

[0101] Exemplary IS110 family IS element sequences are shown in Figure 15 (SEQ ID NOS: 1-348). The nucleotide sequences of the LE, core (if present), transposase, and RE are shown as described above for Figure 15. Additional exemplary IS110 family IS element sequences are shown in SEQ ID NOS: 349-10175. The nucleotide sequences of the LE, core (if present), transposase CDS, RE, and bridge RNA are shown as features in the sequence listing.

[0102] For IS110 elements that contain a core sequence, RE-core-LE refers to the concatenation of the RE, core, and LE nucleotide sequences, a portion of which (e.g., a donor site sequence consisting of LD-core-RD) can be bound by IS110 family transposases described herein (see, e.g., Section C). In some embodiments, RE-core-LE is a sequence similar to that described in Figure 15 (SEQ ID NOS: 1-348) or Figure 17 (SEQ ID NOS: 30354-30356). 30529). In some embodiments, RE-core-LE comprises an LE, core, and RE set forth in SEQ ID NOs: 349-10175 or 30530-40356. The nucleotide sequences of the LE, core (if present), and RE are set forth as features in this sequence listing. In some embodiments, LD-core-RD comprises an LD, core, and RD set forth in Figure 17 (SEQ ID NOs: 30354-30529). The nucleotide sequences of the LD and RD are shown in bold, and the nucleotide sequence of the core sequence is presented in single-underlined, unhighlighted text. In some embodiments, LD-core-RD comprises an LD, core, and RD derived from the LDG and RDG set forth in Figure 15 (SEQ ID NOs: 1-348).

[0103] For IS110 elements containing a core sequence, LF-core-RF refers to the concatenation of the LF, core, and RF nucleotide sequences, a portion of which (e.g., a target site sequence consisting of LT-core-RT) can be bound by IS110 family transposases described herein (see, e.g., Section C). In some embodiments, LF-core-RF comprises the LF, core, and RF set forth in Figure 18 (SEQ ID NOS: 20351-20526). In some embodiments, RF-core-LF comprises the LF, core, and RF set forth in SEQ ID NOS: 20527-30353. The nucleotide sequences of the LF, core (if present), and RF are set forth as features in this sequence listing. In some embodiments, LT-core-RT comprises the LT, core, and RT set forth in Figure 18 (SEQ ID NOS: 20351-20368). The nucleotide sequences of the LT and RT are shown in bold, and the nucleotide sequence of the core sequence is presented in single-underlined, unhighlighted text. In some embodiments, the LT-core-RT comprises LT, core, and RT derived from the LTG and RTG shown in Figure 15 (SEQ ID NOs: 1-348).

[0104] For IS110 elements that do not contain a core sequence, RE-LE refers to the concatenation of the nucleotide sequences of RE and LE, a portion of which (e.g., a donor site sequence consisting of LD-RD) can be bound by IS110 family transposases described herein (see, e.g., Section C). In some embodiments, an RE-LE comprises an LE and RE set forth in Figure 15 (SEQ ID NOs: 1-348) or Figure 17 (SEQ ID NOs: 30354-30529). In some embodiments, an RE-LE comprises an LE and RE set forth in SEQ ID NOs: 349-10175 or 30530-40356. In some embodiments, an LD-RD comprises an LD and RD derived from an LDG and RDG set forth in Figure 15 (SEQ ID NOs: 1-348).

[0105] For IS110 elements that do not contain a core sequence, LF-RF refers to the concatenation of the nucleotide sequences of LF and RF, a portion of which (e.g., a target site sequence consisting of LT-RT) can be bound by IS110 family transposases described herein (see, e.g., Section C). In some embodiments, LF-RF comprises the LF and RF set forth in Figure 18 (SEQ ID NOs: 20351-20526). In some embodiments, LF-RF comprises the LF and RF set forth in SEQ ID NOs: 20527-30353. In some embodiments, LT-RT comprises the LT and RT derived from the LTG and RTG set forth in Figure 15 (SEQ ID NOs: 1-348).

[0106] C. IS110 Family Transposases The IS110 family of transposases encoded within the IS110 transposon was identified by homology searches of the DEDD catalytic domain, a RuvC-like domain. See Example 1. The IS110 family transposases described herein comprise an N-terminal RuvC-like DEDD catalytic domain and a C-terminal transposase domain with two canonical Pfam domains, as shown in Figure 1B. In some embodiments, the N-terminal RuvC-like DEDD catalytic domain and the C-terminal transposase domain The polypeptide sequence between comprises a linker domain comprising a coiled coil.

[0107] Within the IS110 family, transposons can be classified into the IS110 group, which is any insertion sequence (IS) element that encodes the IS110 transposase and contains a 5' non-coding end (LE) that is longer than the 3' non-coding end (RE). See Figure 1A, D, E.

[0108] Within the IS110 family, transposons can be classified into the IS1111 group, which is any insertion sequence (IS) element that encodes the IS110 transposase and typically contains a 3' non-coding end (RE) that is longer than the 5' non-coding end (LE) (see Figure 1A, D, E).

[0109] Exemplary primary amino acid sequences and predicted secondary structures of IS110 family transposases are shown in Figure 16 (SEQ ID NOS: 10176-10523). Additional exemplary primary amino acid sequences of IS110 family transposases are shown in SEQ ID NOS: 10524-20350 and 40357-516430. A transposase domain with a RuvC-like DEDD catalytic domain and two canonical Pfam domains is shown as a feature in this sequence listing. The polypeptide sequence between the N-terminal RuvC-like DEDD catalytic domain and the C-terminal transposase domain contains a linker domain containing a coiled coil.

[0110] In some embodiments, the IS110 family transposase has a nucleotide sequence similar to that of the sequence shown in FIG. 16 (SEQ ID NOs: 10176-10523) or SEQ ID NOs: 10524-20350 or 40357-516430, but not greater than 50%, 51% or greater, 52% or greater, 53% or greater, 54% or greater, 55% or greater, 56% or greater, 57% or greater, 58% or greater, 59% or greater, 60% or greater, 61% or greater, 62% or greater, 63% or greater, 64% or greater, 65% or greater, 66% or greater, 67% or greater, 68% or greater, 69% or greater, 70% or greater, 71% or greater, 72% or greater, 73% or greater, 74% or greater, 75% or greater, 76% or greater, 77% or greater, 78% or greater, 79% or greater, 80% or greater, 81% or greater, 82% or greater, 83% or greater, 84% or greater, 85% or greater, 86% or greater, 87% or greater, 88% or greater, 89% or greater, 90% or greater, 91% or greater, 92% or greater, 93% or greater, 94% or greater, 95% or greater, 96% or greater, 97% or greater, 98% or greater, 99% or greater, 100% or greater, 101% or greater, 102% or greater, 103% or greater, 104% or greater, 105% or greater, 106% or greater, 107% or greater, 108% or greater, 10 %, 69% or more, 70% or more, 71% or more, 72% or more, 73% or more, 74% or more, 75% or more, 76% or more, 77% or more, 78% or more, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, or 100% identical to the amino acid sequence of "protein_IS621" of Figure 16 (SEQ ID NO: 10176).

[0111] Domain motifs and / or regions of IS110 family transposases can be identified by structural similarity, not necessarily by amino acid sequence similarity. In some embodiments, structural similarity is determined by template modeling score (TM score). See Example 15 and Figures 14A-H. In some embodiments, predicted secondary structure is used to identify domain motifs and / or regions of IS110 family transposases. In some embodiments, the secondary structure of a primary amino acid sequence is predicted using standard mkdssp tools or equivalent protein structure prediction software on a tertiary structure file. In some embodiments, the linker domain of an IS110 family transposase comprises a polypeptide sequence between a RuvC-like DEDD catalytic domain and a transposase domain comprising an amino acid sequence predicted to form a coiled coil.

[0112] In some embodiments, an IS110-family transposase comprises a polypeptide that forms a tertiary structure similar to the tertiary structure of IS621 shown in Figure 14C. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is similar to the tertiary structure of IS621 if it has a template modeling score (TM score) of 0.5 or greater. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of IS621 if the template modeling score (TM score) is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater.

[0113] The TM score is defined as shown in Figure 14A, where L target is the length of the amino acid sequence of the target protein, and L common is the number of residues that appear in both the template and target structures. i is the distance between the ith residue pair in the template structure and the target structure, and d o (L target )=1.24 3 √(L_target-15)-1.8 is the distance scale that normalizes the distance (Zhang, Yang, and Jeffrey Skolnick. 2004. "Scoring Function for Automated Assessment of Protein Structure Template Quality." Proteins 57 (4): 702-10). Alternatively, the score can be normalized by the length of the query protein, or the score can be normalized by the average length of the two proteins. The TM score has a value in (0,1], and a cutoff of >0.5 is commonly used to identify proteins with homologous tertiary structures (Zhang, Yang, and Jeffrey Skolnick. 2005. “TM-Align: A Protein Structure Alignment Algorithm Based on the TM-Score.” Nucleic Acids Research 33 (7): 2302-9)。

[0114] In some embodiments, the IS110 family transposase has a nucleotide sequence similar to that of the sequence set forth in FIG. 16 (SEQ ID NOs: 10176-10523) or SEQ ID NOs: 10524-20350 or 40357-516430, but is not specifically limited to a specific sequence. In some embodiments, the IS110 family transposase has a nucleotide sequence similar to that of the sequence set forth in FIG. 16 (SEQ ID NOs: 10176-10523), or SEQ ID NOs: 10524-20350 or 40357-516430, but is not specifically limited to a specific ... or more, 26% or more, 27% or more, 28% or more, 29% or more, 30% or more, 31% or more, 32% or more, 33% or more, 34% or more, 35% or more, 36% or more, 37% or more, 38% or more, 39% or more, 40% or more, 41% or more, 42% or more, 43% or more, 44% or more, 45% or more, 46% or more, 47% or more, 48% or more, 49% or more, 50% or more, 51% or more, 52% or more, 53% or more, 54% or more, 55% or more, 56% or more, 57% or more, 58% or more, 59% or more, 60% or more, 61% or more, 62% or more, 63% or more, 64% or more, 65% or more, 66% or more, 67% or more, 68% or more, 69% or more, 70% or more, 71% or more, 72% or more, 73% or more, 74% or more, 75% or more, 76% or more, 77% or more, 78% or more, 79% or more, 80% 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, or 100% identical, and form a similar tertiary structure as shown in Figure 14C. In some embodiments, the sequence is "protein_IS621" (SEQ ID NO: 10176) in Figure 16. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of IS621 if the template modeling score (TM score) is 0.5 or greater. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of IS621 if the template modeling score (TM score) is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater.

[0115] In certain aspects, described herein are IS110 family transposases comprising means for carrying out a transposase reaction. The means for carrying out the transposase reaction comprises the sequences shown in Figure 16 (SEQ ID NOs: 10176-10523) or SEQ ID NOs: 10524-20350 or 40357-516430.

[0116] In certain aspects, described herein are nucleic acids that encode any of the amino acid sequences of the IS110 family transposases provided herein.

[0117] C.1. RuvC-like DEDD catalytic domain The RuvC-like DEDD catalytic domain refers to a domain of IS110 transposase that is similar to the RuvC Holliday junction resolvase, which is an abundant protein domain found in proteins of diverse functions. The RuvC domain is often found in RNA-guided CRISPR nucleases. CRISPR nucleases with RNA-guided RuvC domains can be associated with transposons, such as CRISPR-associated transposons (CASTs). CRISPR nucleases associated with transposons do not mediate transposition, but rather confer target specificity to the transposome. The IS110 family transposases described herein contain a RuvC-like DEDD catalytic domain.

[0118] In some embodiments, the IS110 family transposase has a sequence similar to that of the RuvC-like DEDD catalytic domain sequence set forth in FIG. 16 (SEQ ID NOs: 10176-10523) or SEQ ID NOs: 10524-20350 or 40357-516430, but is not specifically limited to a sequence similar to that of the RuvC-like DEDD catalytic domain sequence set forth in FIG. 16 (SEQ ID NOs: 10176-10523), or SEQ ID NOs: 10524-20350 or 40357-516430, and is similar to that of the RuvC-like DEDD catalytic domain sequence set forth in FIG. 16 (SEQ ID NOs: 10176-10523), or SEQ ID NOs: 10524-20350 or 40357-516430, but is not specifically limited to a sequence ... or 100% identical to a RuvC-like DEDD catalytic domain sequence of "protein_IS621" in FIG. 16 . In some embodiments, the RuvC-like DEDD catalytic domain sequence is the RuvC-like DEDD catalytic domain of "protein_IS621" in FIG. 16 .

[0119] In some embodiments, the IS110 family transposase comprises a RuvC-like DEDD catalytic domain that forms a tertiary structure similar to that of the IS621 RuvC-like DEDD catalytic domain. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 RuvC-like DEDD catalytic domain if the template modeling score (TM score) of the RuvC-like DEDD catalytic domain is 0.5 or greater. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 RuvC-like DEDD catalytic domain if the template modeling score (TM score) is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater.

[0120] In some embodiments, certain amino acid sequences are considered to form a tertiary structure similar to the IS621 RuvC-like DEDD catalytic domain based on the distance between the alpha carbons of conserved residues within the RuvC-like DEDD catalytic domain. Figure 16 (SEQ ID NOS: 10176-10523) bolds the five most highly conserved amino acids in the RuvC-like DEDD catalytic domain. SEQ ID NOS: 10524-20350 or 40357-516430 bolds the five most highly conserved amino acids in the RuvC-like DEDD catalytic domain. The conserved amino acids are shown as features P1-P5 in this sequence listing. Conserved amino acids in a particular amino acid sequence are identified by primary amino acid sequence alignment. In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the RuvC-like DEDD catalytic domain of IS621 if the distance between the alpha carbon of the first conserved residue and the alpha carbon of the second conserved residue in the amino acid sequence ("D1") is less than 10 angstroms (Å), where the conserved residues in the amino acid sequence are based on primary amino acid sequence alignment with one or more RuvC-like DEDD catalytic domains, such as IS621. In IS621, the first conserved residue is D11 and the second conserved residue is E60. In some embodiments, D1 is between 4 and 10 angstroms (Å). In some embodiments, D1 is between 5 and 7.5 angstroms (Å). In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to the IS621 RuvC-like DEDD catalytic domain if the average distances ("D2") between the alpha carbon of the first conserved residue and the alpha carbon of the third conserved residue, between the alpha carbon of the first conserved residue and the alpha carbon of the fourth conserved residue, and between the alpha carbon of the first conserved residue and the alpha carbon of the fifth conserved residue are less than 10 angstroms (Å), where the conserved residues of the amino acid sequence are based on primary amino acid sequence alignment with one or more RuvC-like DEDD catalytic domains, such as IS621. In IS621, the first conserved residue is D11, the third conserved residue is K100, the fourth conserved residue is D102, and the fifth conserved residue is D105. In some embodiments, D2 is between 5 and 10 angstroms (Å). In some embodiments, D2 is between 7.5 and 10 angstroms (Å).In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to the IS621 RuvC-like DEDD catalytic domain if the average distances ("D3") between the alpha carbon of the second conserved residue and the alpha carbon of the third conserved residue, between the alpha carbon of the second conserved residue and the alpha carbon of the fourth conserved residue, and between the alpha carbon of the second conserved residue and the alpha carbon of the fifth conserved residue are less than 15 angstroms (Å), where the conserved residues of the amino acid sequence are based on primary amino acid sequence alignment with one or more RuvC-like DEDD catalytic domains, such as IS621. In IS621, the second conserved residue is E60, the third conserved residue is K100, the fourth conserved residue is D102, and the fifth conserved residue is D105. In some embodiments, D3 is 10-15 angstroms (Å). In some embodiments, D3 is 13-15 angstroms (Å). In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 RuvC-like DEDD catalytic domain when D1 is less than 10 angstroms (Å), D2 is less than 10 angstroms (Å), and D3 is less than 15 angstroms (Å). In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 RuvC-like DEDD catalytic domain when D1 is 4-10 angstroms (Å), D2 is 5-10 angstroms (Å), and D3 is 10-15 angstroms (Å). In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 RuvC-like DEDD catalytic domain when D1 is 5-7.5 angstroms (Å), D2 is 7.5-10 angstroms (Å), and D3 is 13-15 angstroms (Å).

[0121] In some embodiments, the IS110 family transposase has a sequence similar to that of the RuvC-like DEDD catalytic domain sequence shown in FIG. 16 (SEQ ID NOs: 10176-10523) or SEQ ID NOs: 10524-20350 or 40357-516430, but is not specifically limited to a sequence similar to that of the IS110 family transposase. In some embodiments, the IS110 family transposase has a sequence similar to that of the RuvC-like DEDD catalytic domain sequence shown in FIG. 16 (SEQ ID NOs: 10176-10523), or SEQ ID NOs: 10524-20350 or 40357-516430, but is not specifically limited to a sequence similar to that of the IS110 family transposase. or more, 26% or more, 27% or more, 28% or more, 29% or more, 30% or more, 31% or more, 32% or more, 33% or more, 34% or more, 35% or more, 36% or more, 37% or more, 38% or more, 39% or more, 40% or more, 41% or more, 42% or more, 43% or more, 44% or more, 45% or more, 46% or more, 47% or more, 48% or more, 49% or more, 50% or more, 51% or more, 52% or more, 53% or more, 54% or more, 55% or more, 56% or more, 57% or more, 58% or more, 59% or more, 60% or more, 61% or more, 62% or more, 63% or more, 64% or more, 65% or more, 66% or more, 67% or more, 68% or more, 69% or more, 70% or more, 71% or more, 72% or more, 73% or more, 74% or more, 75% or more, 76% or more, 77% or more, 78% or more, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more , 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, or 100% identical, and form a tertiary structure similar to the IS621 RuvC-like DEDD catalytic domain. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to the IS621 RuvC-like DEDD catalytic domain if the template modeling score (TM score) of the RuvC-like DEDD catalytic domain is 0.5 or greater. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to the IS621 RuvC-like DEDD catalytic domain if the template modeling score (TM score) is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater. In some embodiments, the RuvC-like DEDD catalytic domain sequence provided is the RuvC-like DEDD catalytic domain of "protein_IS621" in Figure 16 (SEQ ID NO: 10176).

[0122] In some embodiments, RuvC-like DEDD catalytic domains can be identified using statistical models that annotate protein domains, such as the Pfam Profile Hidden Markov Model (pHMM). In IS110 family transposases, RuvC-like DEDD catalytic domains are often recognized by the Pfam Profile Hidden Markov Model (pHMM) PF01548, abbreviated DEDD_Tnp_IS110.

[0123] In some embodiments, the IS110 family transposase is Dx(43)-Ex(39)-Kx(1)-Dx(2)-D (SEQ ID NO: 795142), Dx(42)-Ex(34)-Kx(1)-Dx(2)-D (SEQ ID NO: 795143), [DE]-x(38,63)-[EACDGQVIPS]-x(30,53)-[KSQIVRHLTMA]-x(1)-[DNE]-x(2)-[DEASCM], [DE]-x(41,59)-[EYALGHCVFITMS]-x(30,45)-[KRMQ NH]-x(1)-[DN]-x(2)-[DAS], GIDVS (SEQ ID NO: 795144), GLDVH (SEQ ID NO: 795145), [GAS][ILVDFMACWTHGN][DE][VTIWPLAFCSMR][SAHGCD], [GAS][LIMVCFA]D[VLIFAYQTHMDCWRGS][HASGD], MEATG (SEQ ID NO: 795146), MEACG (SEQ ID NO: 795147), [MLVIFCYAGTWSHPR]-x(0,2)-[EGACQDVIMP][ASGPHID RYLNQCTVFMEK][TSECPAYGIVDNFKLMHR]-x(0,1)-[GASTRDQVNLWEKYHICMP], [MYVCLIASFTGQEHW]-x(0,2)-[EYQGVACFLHITMSDKW]-x(0,2)-[AEVYSFMGTICNPLDQ]-x(0,2)-[CGATESMVLINPFDQW]-x(0,2)-[GPASTLCYINRMVFEWHQDK], DRIDA (SEQ ID NO: 795148), DRRDA (SEQ ID NO: 795149), [DNE][ [RPAKVTSQEFGIDLHMNYWC][ILKVAGTNRSFQPHDEMCWY][DAESCM][ACSPLTGV], or a RuvC-like DEDD catalytic domain containing the motif [DN][RAKEYDQGFVPTMSLHWNIC][RALNIVKHTQDSEMGWFYCP][DSA][ASTGCVLI], where the motif is of the general Prosite format, where x is any amino acid, x(n) represents n any amino acids, and x(n,m) represents n to m any amino acids. In some embodiments, Ru The vC-like DEDD catalytic domain comprises the motif Dx(43)-Ex(39)-Kx(1)-Dx(2)-D (SEQ ID NO: 795142), Dx(42)-Ex(34)-Kx(1)-Dx(2)-D (SEQ ID NO: 795143), GIDVS (SEQ ID NO: 795144), GLDVH (SEQ ID NO: 795145), MEATG (SEQ ID NO: 795146), MEACG (SEQ ID NO: 795147), DRIDA (SEQ ID NO: 795148), or DRRDA (SEQ ID NO: 795149). In some embodiments, RuvC-like DEDD catalytic domains comprising the above motifs form a tertiary structure similar to the RuvC-like DEDD catalytic domain of IS621. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 RuvC-like DEDD catalytic domain if the template modeling score (TM score) of the RuvC-like DEDD catalytic domain is 0.5 or greater. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 RuvC-like DEDD catalytic domain if the template modeling score (TM score) is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater.

[0124] In some embodiments, the IS110 family transposase has a sequence similar to that of the RuvC-like DEDD catalytic domain sequence of FIG. 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350 or 40357-516430, but is 50% or more, 51% or more, 52% or more, 53% or more, 54% or more, 55% or more, 56% or more, 57% or more, 58% or more, 59% or more, 60% or more, 61% or more, 62% or more, 63% or more, 64% or more, 65% or more, 66% or more, 67% or more, 68% or more, 69% or more, 70% or more, 71% or more, or 72% or more of the sequence similar to that of the RuvC-like DEDD catalytic domain sequence of FIG. 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350 or 40357-516430. 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, or 100% identical to the RuvC-like DEDD catalytic domain of IS621, and comprising a motif in any of the preceding paragraphs or Figures 21-28. In some embodiments, the amino acid sequence forms a tertiary structure similar to the RuvC-like DEDD catalytic domain of IS621. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 RuvC-like DEDD catalytic domain if the template modeling score (TM score) of the RuvC-like DEDD catalytic domain is 0.5 or greater. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 RuvC-like DEDD catalytic domain if the template modeling score (TM score) is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater.

[0125] In some embodiments, an IS110 transposase belonging to the IS110 group comprises a RuvC-like DEDD catalytic domain comprising the domain motif shown in Figure 21, where the motif is of the general Prosite format, where x is any amino acid, x(n) represents any n amino acids, and x(n,m) represents any n to m amino acids. In some embodiments, an IS110 transposase belonging to the IS110 group comprises a RuvC-like DEDD catalytic domain comprising one or more of the RuvC-like DEDD catalytic domain motifs shown in Figures 22-24, where the motif is of the general Prosite format, where x is any amino acid, x(n) represents any n amino acids, and x(n,m) represents any n to m amino acids. In some embodiments, an IS110 transposase belonging to the IS110 group comprises a RuvC-like DEDD catalytic domain wherein the "D" region of the canonical DEDD catalytic motif comprises the domain motif shown in Figure 22, where the motif is a general Prosite format, where x is any amino acid, x(n) represents any n amino acids, and x(n,m) represents any n to m amino acids. In some embodiments, IS110 transposases belonging to the IS110 group comprise a RuvC-like DEDD catalytic domain in which the "E" region of the canonical DEDD catalytic motif comprises the domain motif shown in Figure 23, where the motif is a general Prosite format, where x is any amino acid, x(n) represents any n amino acids, and x(n,m) represents any n to m amino acids. In some embodiments, IS110 transposases belonging to the IS110 group comprise a RuvC-like DEDD catalytic domain in which the "DD" region of the canonical DEDD catalytic motif comprises the domain motif shown in Figure 24, where the motif is a general Prosite format, where x is any amino acid, x(n) represents any n amino acids, and x(n,m) represents any n to m amino acids. In some embodiments, an IS110 transposase belonging to the IS110 group comprises a RuvC-like DEDD catalytic domain in which the "D" region of the canonical DEDD catalytic motif comprises the domain motif shown in Figure 22, the "E" region of the canonical DEDD catalytic motif comprises the domain motif shown in Figure 23, and the "DD" region of the canonical DEDD catalytic motif comprises the domain motif shown in Figure 24, wherein the motifs are of the general Prosite format, where x is any amino acid, x(n) represents n any amino acids, and x(n,m) represents n to m any amino acids.

[0126] In some embodiments, IS1110 transposases belonging to the IS1111 group comprise a RuvC-like DEDD catalytic domain comprising the domain motif shown in Figure 25, where the motif is in the general Prosite format, where x is any amino acid, x(n) represents any n amino acids, and x(n,m) represents any n to m amino acids. In some embodiments, IS1110 transposases belonging to the IS1111 group comprise a RuvC-like DEDD catalytic domain comprising one or more of the RuvC-like DEDD catalytic domain motifs shown in Figures 26-28, where the motif is in the general Prosite format, where x is any amino acid, x(n) represents any n amino acids, and x(n,m) represents any n to m amino acids. In some embodiments, IS1110 transposases belonging to the IS1111 group comprise a RuvC-like DEDD catalytic domain in which the "D" region of the canonical DEDD catalytic motif comprises the domain motif shown in Figure 26, where the motif is of the general Prosite format, where x is any amino acid, x(n) represents any n amino acids, and x(n,m) represents any n to m amino acids. In some embodiments, IS1111 group transposases belong to the RuvC-like DEDD catalytic domain in which the "E" region of the canonical DEDD catalytic motif comprises the domain motif shown in Figure 27, where the motif is of the general Prosite format, where x is any amino acid, x(n) represents any n amino acids, and x(n,m) represents any n to m amino acids. In some embodiments, an IS110 transposase belonging to the IS1111 group comprises a RuvC-like DEDD catalytic domain in which the "DD" region of the canonical DEDD catalytic motif comprises the domain motif shown in Figure 28, where the motif is of the general Prosite format, where x is any amino acid, x(n) represents n any amino acids, and x(n,m) represents n to m any amino acids.In some embodiments, an IS110 transposase belonging to the IS1111 group comprises a RuvC-like DEDD catalytic domain in which the "D" region of the canonical DEDD catalytic motif comprises the domain motif shown in Figure 26, the "E" region of the canonical DEDD catalytic motif comprises the domain motif shown in Figure 27, and the "DD" region of the canonical DEDD catalytic motif comprises the domain motif shown in Figure 28, wherein the motifs are of the general Prosite format, where x is any amino acid, x(n) represents n any amino acids, and x(n,m) represents n to m any amino acids.

[0127] The present invention provides a method for generating IS110 transposase chimeras with advantageous functions. Main swapping is also contemplated. In some embodiments, as described in Farruggio et al., 2014, swapping of a RuvC-like DEDD catalytic domain (e.g., any RuvC-like DEDD catalytic domain shown in FIG. 16 (SEQ ID NOs: 10176-10523) or SEQ ID NOs: 10524-20350 or 40357-516430) with a different RuvC-like DEDD catalytic domain (e.g., any other RuvC-like DEDD catalytic domain shown in FIG. 16 (SEQ ID NOs: 10176-10523) or SEQ ID NOs: 10524-20350 or 40357-516430) that results in advantageous properties is also contemplated. For example, swapping one RuvC-like DEDD catalytic domain for another RuvC-like DEDD catalytic domain may allow for higher affinity for the bridge RNA or increased transfer efficiency.

[0128] C.2. Transposase Domain The IS110 family transposases described herein comprise a transposase domain.

[0129] In some embodiments, transposase domains can be identified using statistical models that annotate protein domains, such as the Pfam Profile Hidden Markov Model (pHMM). In IS110 family transposases, the transposase domain is often recognized by the Pfam Profile Hidden Markov Model (pHMM) PF02371, abbreviated as Transposase_20.

[0130] In some embodiments, the IS110 family transposase has a sequence similar to or different from the sequence of the transposase domain provided in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350 or 40357-516430, but is not specifically limited to a sequence similar to or different from the sequence of the transposase domain provided in Figure 16 (SEQ ID NOS: 10176-10523), or SEQ ID NOS: 10524-20350 or 40357-516430, and is similar to or different from the sequence of the transposase domain provided in Figure 16 (SEQ ID NOS: 10176-10523), or SEQ ID NOS: 10524-20350 or 40357-516430, but is not specifically limited to a sequence similar to ... 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, or 100% identical to the transposase domain sequence of "protein_IS621" in Figure 16 (SEQ ID NO: 10176).

[0131] In some embodiments, the IS110 family transposase comprises a transposase domain that forms a tertiary structure similar to that of the IS621 transposase domain. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 transposase domain if the template modeling score (TM score) of the transposase domain is 0.5 or greater. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 transposase domain if the template modeling score (TM score) is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater.

[0132] In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to the transposase domain of IS621 based on the distance between the alpha carbons of conserved residues in the transposase domain. 3) indicates in bold the five most highly conserved amino acids in the transposase domain. SEQ ID NOS: 10524-20350 or 40357-516430 indicate the five most highly conserved amino acids in the transposase domain as Features P1-P5 in the Sequence Listing. Conserved amino acids in a particular amino acid sequence are identified by alignment of the primary amino acid sequence. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to the IS621 transposase domain if the average distance ("D1") between the alpha carbon of the first conserved residue and the alpha carbon of the second conserved residue, and between the alpha carbon of the first conserved residue and the alpha carbon of the fifth conserved residue of the amino acid sequence is less than 25 angstroms (Å), where the conserved residues of the amino acid sequence are based on alignment of the primary amino acid sequence with one or more transposase domains, such as IS621. In IS621, the first conserved residue is G203, the second conserved residue is G233, and the fifth conserved residue is G255. In some embodiments, D1 is 15-25 angstroms (Å). In some embodiments, D1 is 17-23 angstroms (Å). In some embodiments, a particular amino acid sequence is considered to form a similar tertiary structure to the transposase domain of IS621 if the average distances ("D2") between the alpha carbon of the second conserved residue and the alpha carbon of the third conserved residue, between the alpha carbon of the second conserved residue and the alpha carbon of the fourth conserved residue, between the alpha carbon of the fifth conserved residue and the alpha carbon of the third conserved residue, and between the alpha carbon of the fifth conserved residue and the alpha carbon of the fourth conserved residue are less than 25 angstroms (Å), where the conserved residues of the amino acid sequence are based on an alignment of the primary amino acid sequence with one or more transposase domains, such as IS621. In IS621, the second conserved residue is G233, the third conserved residue is S241, the fourth conserved residue is G242, and the fifth conserved residue is G255. In some embodiments, D2 is 20-25 angstroms (Å). In some embodiments, D2 is 22-24 angstroms (Å).In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 transposase if the distance between the alpha carbon of the second conserved residue and the alpha carbon of the fifth conserved residue ("D3") is less than 15 angstroms (Å), where the conserved residues of the amino acid sequence are determined by alignment of the primary amino acid sequence with one or more transposase domains, such as IS621. In IS621, the second conserved residue is G233 and the fifth conserved residue is G255. In some embodiments, D3 is between 5 and 15 angstroms (Å). In some embodiments, D3 is between 7 and 12 angstroms (Å). In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 transposase domain if D1 is less than 25 angstroms (Å), D2 is less than 25 angstroms (Å), and D3 is less than 15 angstroms (Å). In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to the IS621 transposase domain when D1 is 15-25 angstroms (Å), D2 is 20-25 angstroms (Å), and D3 is 5-15 angstroms (Å). In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to the IS621 transposase domain when D1 is 17-23 angstroms (Å), D2 is 22-24 angstroms (Å), and D3 is 7-12 angstroms (Å).

[0133] In some embodiments, the IS110 family transposase has a sequence similar to that of the transposase domain sequence shown in FIG. 16 (SEQ ID NOs: 10176-10523), or SEQ ID NOs: 10524-20350, or 40357-516430, but is not specifically restricted to the sequence shown in FIG. 16 (SEQ ID NOs: 10176-10523), or SEQ ID NOs: 10524-20350, or 40357-516430, and is 15% or more, 16% or more, 17% or more, 18% or more, 19% or more, 20% or more, 21% or more, 22% or more, 23% or more, 24% or more, 25% or more, 26% or more, 27% or more, 28% or more, 29% or more, 30% or more, 31% or more, 32% or more, 33% or more, 34% or more, 35% or more, 36% or more, 37% or more, 38% or more, 39% or more, 40% or more, 41% or more, 42% or more Above, 43% or more, 44% or more, 45% or more, 46% or more, 47% or more, 48% or more, 49% or more, 50% or more, 51% or more, 52% or more, 53% or more, 54% or more, 55% or more, 56% or more, 57% or more, 58% or more, 59% or more, 60% or more, 61% or more, 62% or more, 63% or more, 64% or more, 65% or more, 66% or more, 67% or more, 68% or more, 69% or more, 70% or more, 71% or more, 72% or more, 73% or more, 74% or more, 75% or more, 76% or more, 77% or more , 78% or more, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, or 100% identical, and the amino acid sequence forms a tertiary structure similar to the IS621 transposase domain. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to the IS621 transposase domain if the template modeling score (TM score) of the transposase domain is 0.5 or greater. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to the transposase domain of IS621 if the template modeling score (TM score) is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater. In some embodiments, the transposase domain sequence provided is the transposase domain of "protein_IS621" (SEQ ID NO: 10176) in Figure 16.

[0134] In some embodiments, the IS110-family transposase comprises a RuvC-like DEDD catalytic domain as described in the preceding paragraph (see Section C.1.) and further comprises a sequence identity that is 50% or more, 51% or more, 52% or more, 53% or more, 54% or more, 55% or more, 56% or more, 57% or more, 58% or more, 59% or more, 60% or more, 61% or more, 62% or more, 63% or more, 64% or more, or 65% or more of the transposase domain sequence shown in Figure 16 (SEQ ID NOS: 10176-10523) or SEQ ID NOS: 10524-20350 or 40357-516430. %, 65% or more, 66% or more, 67% or more, 68% or more, 69% or more, 70% or more, 71% or more, 72% or more, 73% or more, 74% or more, 75% or more, 76% or more, 77% or more, 78% or more, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, or 100% identical. In some embodiments, the amino acid sequence forms a tertiary structure similar to the transposase domain of IS621. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 transposase domain if the template modeling score (TM score) of the transposase domain is 0.5 or greater. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 transposase domain if the template modeling score (TM score) is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater. In some embodiments, the transposase domain sequence provided is the transposase domain of "protein_IS621" (SEQ ID NO: 10176) in Figure 16.

[0135] In some embodiments, the IS110-family transposase comprises a RuvC-like DEDD catalytic domain as described in the previous paragraph (see Section C.1.) and further comprises a transposase having a tertiary structure whose amino acid sequence is similar to the transposase domain of IS621. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 transposase domain if the template modeling score (TM score) of the transposase domain is 0.5 or greater. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 transposase domain if the template modeling score (TM score) is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater. In some embodiments, the transposase domain sequence provided is the transposase domain of "protein_IS621" (SEQ ID NO: 10176) in Figure 16. In some embodiments, the amino acid sequence forms a tertiary structure similar to that of the IS621 RuvC-like DEDD catalytic domain. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 RuvC-like DEDD catalytic domain if the RuvC-like DEDD catalytic domain has a template modeling score (TM score) of 0.5 or greater. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 RuvC-like DEDD catalytic domain if the template modeling score (TM score) is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater. In some embodiments, the RuvC-like DEDD catalytic domain sequence provided is the RuvC-like DEDD catalytic domain of "protein_IS621" (SEQ ID NO: 10176) in Figure 16.In some embodiments, the transposase domain has a sequence identity that is 15% or more, 16% or more, 17% or more, 18% or more, 19% or more, 20% or more, 21% or more, 22% or more, 23% or more, 24% or more, 25% or more, 26% or more, 27% or more, 28% or more, 29% or more, 30% or more, 31% or more, 32% or more, 33% or more, 34% or more, 35% or more, 36% or more, 37% or more, 38% or more, 39% or more, 40% or more, 41% or more, 42% or more, 43% or more, 44% or more, 45% or more, 46% or more, 47% or more, 48% or more, 49% or more, 50% or more, or 60% or more of the transposase domain sequence shown in FIG. 16 (SEQ ID NOs: 10176-10523) or SEQ ID NOs: 10524-20350 or 40357-516430. % or more, 51% or more, 52% or more, 53% or more, 54% or more, 55% or more, 56% or more, 57% or more, 58% or more, 59% or more, 60% or more, 61% or more, 62% or more, 63% or more, 64% or more, 65% or more, 66% or more, 67% or more, 68% or more, 69% or more, 70% or more, 71% or more, 72% or more, 73% or more, 74% or more, 75% or more, 76% or more , 77% or more, 78% or more, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, or 100% identical.

[0136] In some embodiments, the IS110 family transposase has a sequence identity of 50% or more, 51% or more, 52% or more, 53% or more, 54% or more, 55% or more, 56% or more, 57% or more, 58% or more, 59% or more, 60% or more, 61% or more, 62% or more, 63% or more, 64% or more, 65% or more, 66% or more, 67% or more, 68% or more, 69% or more, 70% or more, 71% or more, 72% or more, 73% or more, 74% or more, 75% or more, 76% or more, 77% or more, 78% or more, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, 100% or more, 101% or more, 102% or more, 103% or more, 104% or more, 105% or more, 106% or more, 107% or more, 108% or more, 109% or more, 110% or more, 111% or more, 112% or more, 113% or more, 114% or more, 115% or more, 116% or more, 117% or more a Ruv comprising an amino acid sequence that is 7% or more, 68% or more, 69% or more, 70% or more, 71% or more, 72% or more, 73% or more, 74% or more, 75% or more, 76% or more, 77% or more, 78% or more, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, or 100% identical to the Ruv 16 (SEQ ID NOs: 10176 to 10523), or SEQ ID NOs: 10524 to 20350, or 40357 to 516430), and further comprising 50% or more, 51% or more, 52% or more, 53% or more, 54% or more, 55% or more, 56% or more, 57% or more, 58% or more, 59% or more, 60% or more, 61% or more, 62% or more, 63% or more, 64% or more, 65% or more, 66% or more, 67% or more, 68% or more, 69% or more , 70% or more, 71% or more, 72% or more, 73% or more, 74% or more, 75% or more, 76% or more, 77% or more, 78% or more, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, or 100% identical. In some embodiments, the amino acid sequence forms a tertiary structure similar to the RuvC-like DEDD catalytic domain of IS621. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 RuvC-like DEDD catalytic domain if the RuvC-like DEDD catalytic domain has a template modeling score (TM score) of 0.5 or greater. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 RuvC-like DEDD catalytic domain if the template modeling score (TM score) is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater. In some embodiments, the RuvC-like DEDD catalytic domain sequence provided is the RuvC-like DEDD catalytic domain of "protein_IS621" (SEQ ID NO: 10176) in Figure 16. In some embodiments, the amino acid sequence forms a tertiary structure similar to that of the IS621 transposase domain. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software.In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 transposase domain if the template modeling score (TM score) of the transposase domain is 0.5 or greater. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 transposase domain if the template modeling score (TM score) is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater. In some embodiments, the transposase domain sequence provided is the transposase domain of "protein_IS621" (SEQ ID NO: 10176) in Figure 16.

[0137] In some embodiments, the IS110 family transposase comprises the motif Gx(28)-Gx(7)-SG-x(10)-G (SEQ ID NO: 795150), Gx(29)-Gx(7)-SG-x(11)-G (SEQ ID NO: 795151), IPGIG (SEQ ID NO: 795152), IPGVG (SEQ ID NO: 795153), AGLAP (SEQ ID NO: 795154), LGLVP (SEQ ID NO: 795155), [GSRAVECPYQNHFDKIT]-x(26,39)-[GKARCTNSEQHMD]-x(7,8)-[STRP][GASDNQCVER]-x(8,18)-[GACVYTSKRFNHIQDE], [GASFCRTNQKYDEV]-x(25,31)-[GRAL KMSQCED]-x(7)-[STRAP][GNADSVCHER]-x(10,17)-[GSVRCAYKHQTFINED], [IVLAMFQTERCHSKYPWNG][PKDTFYERVSHQANGILCMW][GRASTCEPKVYQNHFDI]-x(0,1)-[IVAFLMCSWTYGKHNPRE]-x(0,2) -[GSDANKQERTMHCP], [IVLHAMFQTCYEPRSGW][PRKESADYTHVNGQICWLMF][GSFAQCYTRKNWLDEVH][VIFLCMYATPSGW]-x(0,2)-[GADSQNKRTEPWHC],[AVLNCSTIGMFYWDHEPQR][GRTKNACSEHQMD][LVTI [MAYSFCQRWKNPGH][ACDTSVNERHFWIYQGKLMP][PVLAWISTCGNMR], or [LAVIFTCSWMGYPRQNKE][GRLAMCKEQDS][LIVTCMKSAFRPQG][VTAICRNDGSQMHLEPYKF][PGKASVIDRLTENQCW], and the motif is of the general Prosite format, where x is any amino acid, x(n) represents n any amino acids, and x(n,m) represents n to m any amino acids. In some embodiments, the transposase domain comprises the motif Gx(28)-Gx(7)-SG-x(10)-G (SEQ ID NO: 795150), Gx(29)-Gx(7)-SG-x(11)-G (SEQ ID NO: 795151), IPGIG (SEQ ID NO: 795152), IPGVG (SEQ ID NO: 795153), AGLAP (SEQ ID NO: 795154), or LGLVP (SEQ ID NO: 795155). In some embodiments, the amino acid sequence forms a tertiary structure similar to the transposase domain of IS621. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to the transposase domain of IS621 if the template modeling score (TM score) of the transposase domain is 0.5 or greater. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the transposase domain of IS621 if the template modeling score (TM score) is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater.

[0138] In some embodiments, the IS110 family transposase has a sequence similar to that of the transposase domain sequence of FIG. 16 (SEQ ID NOs: 10176-10523) or SEQ ID NOs: 10524-20350 or 40357-516430, but is 50% or more, 51% or more, 52% or more, 53% or more, 54% or more, 55% or more, 56% or more, 57% or more, 58% or more, 59% or more, 60% or more, 61% or more, 62% or more, 63% or more, 64% or more, 65% or more, 66% or more, 67% or more, 68% or more, 69% or more, 70% or more, 71% or more, or 72% or more of the sequence similar to that of the transposase domain sequence of FIG. 16 (SEQ ID NOs: 10176-10523) or SEQ ID NOs: 10524-20350 or 40357-516430. 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, or 100% identical to the IS621 transposase domain, and comprising a motif in the preceding paragraph or any of Figures 29-34. In some embodiments, the amino acid sequence forms a tertiary structure similar to the IS621 transposase domain. In some embodiments, the tertiary structure is determined using AlphaFold or similar protein structure prediction software. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 transposase domain if the template modeling score (TM score) of the transposase domain is 0.5 or greater. In some embodiments, a particular amino acid sequence is considered to form a tertiary structure similar to that of the IS621 transposase domain if the template modeling score (TM score) is 0.5 or greater, 0.6 or greater, 0.7 or greater, 0.8 or greater, or 0.9 or greater.

[0139] In some embodiments, IS110 transposases belonging to the IS110 group comprise a transposase domain comprising a domain motif shown in Figure 29, where the motif is of the general Prosite format, where x is any amino acid, x(n) represents any n amino acids, and x(n,m) represents any n to m amino acids. In some embodiments, IS110 transposases belonging to the IS110 group comprise a transposase domain comprising one or more transposase domain motifs shown in Figures 30-31, where the motif is of the general Prosite format, where x is any amino acid. x(n) represents any n amino acids, and x(n,m) represents any n to m amino acids.

[0140] In some embodiments, IS1110 transposases belonging to the IS1111 group comprise a transposase domain comprising a domain motif shown in Figure 32, where the motif is of the general Prosite format, where x is any amino acid, x(n) represents any n amino acids, and x(n,m) represents any n to m amino acids. In some embodiments, IS1110 transposases belonging to the IS1111 group comprise a transposase domain comprising one or more transposase domain motifs shown in Figures 33-34, where the motif is of the general Prosite format, where x is any amino acid, x(n) represents any n amino acids, and x(n,m) represents any n to m amino acids.

[0141] C.3. Linker Domains The IS110 family transposases described herein comprise a linker domain between the RuvC-like DEDD catalytic domain and the transposase domain. In some embodiments, the linker domain comprises a coiled-coil.

[0142] In some embodiments, the IS110 family transposase has a sequence identity of 50% or more, 51% or more, 52% or more, 53% or more, 54% or more, 55% or more, 56% or more, 57% or more, 58% or more, 59% or more, 60% or more, 61% or more, 62% or more, 63% or more, 64% or more, 65% or more, 66% or more, 67% or more, 68% or more, 69% or more, 70% or more, 71% or more, 72% or more, 73% or more, 74% or more, 75% or more, 76% or more, 77% or more, 78% or more, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, 100% or more, 101% or more, 102% or more, 103% or more, 104% or more, 105% or more, 106% or more, 107% or more, 108% or more, 109% or more, 110% or more, 111% or more, 112% or more, 113% or more, 114% or more, 115% or more, 116% or more, 117% or more The linker domain comprises an amino acid sequence that is 8% or more, 69% or more, 70% or more, 71% or more, 72% or more, 73% or more, 74% or more, 75% or more, 76% or more, 77% or more, 78% or more, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, or 100% identical to the amino acid sequence of the target sequence.

[0143] C.4. IS110 transposase In certain aspects, the invention provides IS110 family transposases comprising an N-terminal RuvC-like DEDD catalytic domain comprising the RuvC-like DEDD catalytic domain sequence set forth in Figure 16 (SEQ ID NOs: 10176-10523), or SEQ ID NOs: 10524-20350, or 40357-516430, a linker domain comprising a coiled-coil, and a C-terminal transposase domain comprising the transposase domain sequence set forth in Figure 16 (SEQ ID NOs: 10176-10523), or SEQ ID NOs: 10524-20350, or 40357-516430. In some embodiments, the linker domain comprises the linker domain sequence set forth in Figure 16 (SEQ ID NOs: 10176-10523), or SEQ ID NOs: 10524-20350, or 40357-516430. In some embodiments, the IS110 family transposase comprises "protein_IS621" of Figure 16 (SEQ ID NO: 10176).

[0144] In a specific aspect, the present invention provides an IS110 fragment comprising an N-terminal RuvC-like DEDD catalytic domain consisting of the RuvC-like DEDD catalytic domain sequence shown in Figure 16 (SEQ ID NOs: 10176-10523) or SEQ ID NOs: 10524-20350 or 40357-516430, a linker domain comprising a coiled-coil, and a C-terminal transposase domain consisting of the transposase domain sequence shown in Figure 16 (SEQ ID NOs: 10176-10523) or SEQ ID NOs: 10524-20350 or 40357-516430. In some embodiments, the linker domain consists of the linker domain sequence shown in Figure 16 (SEQ ID NOs: 10176-10523), or SEQ ID NOs: 10524-20350, or 40357-516430. In some embodiments, the IS110-family transposase consists of "protein_IS621" in Figure 16 (SEQ ID NO: 10176).

[0145] In certain aspects, the present invention provides a RuvC-like DEDD catalytic domain comprising any one of the motifs or sequences of Figures 21 to 28 and having a sequence similar to or different from the RuvC-like DEDD catalytic domain sequence shown in Figure 16 (SEQ ID NOs: 10176 to 10523) or SEQ ID NOs: 10524 to 20350 or 40357 to 516430, and having a sequence similar to or different from the RuvC-like DEDD catalytic domain sequence shown in Figure 16 (SEQ ID NOs: 10176 to 10523), or SEQ ID NOs: 10524 to 20350 or 40357 to 516430, and having a sequence similar to or different from the RuvC-like DEDD catalytic domain sequence shown in Figure 16 (SEQ ID NOs: 10176 to 10523), and having a sequence similar to or different from the RuvC-like DEDD catalytic domain sequence shown in Figure 16 (SEQ ID NOs: 10524 to 20350), or SEQ ID NOs: 10524 to 20350, ...524 to 20350 a RuvC-like DEDD catalytic domain that is 69% or more, 70% or more, 71% or more, 72% or more, 73% or more, 74% or more, 75% or more, 76% or more, 77% or more, 78% or more, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, or 100% identical to a RuvC-like DEDD catalytic domain; and a linker domain comprising a coiled coil. 30 to 34, and 50% or more, 51% or more, 52% or more, 53% or more, 54% or more, 55% or more, 56% or more, 57% or more, 58% or more, 59% or more, 60% or more, 61% or more, 62% or more, 63% or more, 64% or more, 65% or more, 66% or more, 67% or more, 68% or more, 69% or more, 70% or more, 71% or more of the transposase domain sequence shown in FIG. 16 (SEQ ID NOs: 10176 to 10523) or SEQ ID NOs: 10524 to 20350 or 40357 to 516430 and a transposase domain that is, from N-terminus to C-terminus, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the IS110 family transposase.In some embodiments, the linker domain has a sequence similar to that of the linker domain sequence shown in FIG. 16 (SEQ ID NOs: 10176 to 10523), or SEQ ID NOs: 10524 to 20350, or 40357 to 516430, but is not particularly limited thereto. ... or more, 69% or more, 70% or more, 71% or more, 72% or more, 73% or more, 74% or more, 75% or more, 76% or more, 77% or more, 78% or more, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, or 100% identical.

[0146] In certain aspects, described herein are amino acid sequences of IS110 family transposases that further comprise a nuclear localization signal (NLS). In some embodiments, the NLS is encoded at the N-terminus of the amino acid sequence of the IS110 family transposase. In some embodiments, the NLS is encoded at the C-terminus of the amino acid sequence of the IS110 family transposase. In some embodiments, the amino acid sequence of the IS110 family transposase further comprises multiple NLSs. In some embodiments, the amino acid sequence of the IS110 family transposase further comprises two NLSs. In some embodiments, the amino acid sequence of the IS110 family transposase is , and three additional NLSs. In some embodiments, the amino acid sequence of the IS110 family transposase comprises more than three additional NLSs. In some embodiments, the amino acid sequence of the IS110 family transposase comprises three additional NLSs at the N-terminus of the amino acid sequence of the IS110 family transposase. In some embodiments, the amino acid sequence of the IS110 family transposase comprises three additional NLSs at the C-terminus of the amino acid sequence of the IS110 family transposase. Any NLS known in the art can be used. In some embodiments, the NLS is an SV40 NLS.

[0147] Without wishing to be bound by theory, the observed transposition reaction of IS110 resembles that of conservative site-specific recombinase (CSSR). Two DNA sequences are specifically recognized by the transposition machinery. A bridge RNA recognizes both the donor site sequence and the target site sequence. The transposase binds to the bridge RNA, and in some embodiments, to the subterminal inverted repeat (STIR) subsequence of the donor site. Excision of the IS110 transposon element is traceless, and the original target sequence is reconstituted to its original pre-insertion sequence.

[0148] In some embodiments, the present invention can be complemented with additional IS110 orthologs with improved efficiency and / or specificity using extensive bioinformatics searches. IS110 systems with improved efficiency, such as in mammalian cells, can be identified by prioritizing candidates present in organisms that naturally grow at physiological temperatures. For example, candidates naturally present in human gut microbes can be prioritized for experimental characterization in human cells. Alternatively, IS110 systems found in the genomes of extremophile microorganisms can be prioritized for experimental characterization due to their potential utility as in vitro temperature-responsive molecular tools. In another embodiment, IS110 systems with high specificity can be prioritized based on predicted properties of the bridge RNA sequence, as further described below (see Section D).

[0149] The present invention also contemplates the use of the IS110 system with cofactors and / or additional domains that may confer new functions to the IS110 system, increase its specificity, and / or increase its efficiency, which may be present within and / or adjacent to the IS110 element.

[0150] In some embodiments, the present invention contemplates an approach to engineering the IS110 system. Scientists have demonstrated that it is possible to infer the identity of ancestral proteins using phylogenetic analysis (Alonso-Lerma, B., Jabalera, Y., Samperio, S. et al., Evolution of CRISPR-associated endonucleases as inferred from resurrected proteins., Nat Microbiol 8, 77-90 (2023)). While these inferred ancestral proteins may not currently exist in nature, they represent the common ancestor of existing protein clades. These ancestral proteins can then be experimentally synthesized and tested to determine whether they possess favorable properties. In one embodiment, IS110 clades with particular long or specific bridge RNA target and / or donor binding loops can be analyzed using this method to further identify ancestral proteins capable of binding to bridge RNA molecules with favorable properties.

[0151] In one embodiment, the present invention can be complemented with additional IS110 orthologs using extensive bioinformatics searches. Some IS110 transposases can be fused with other domains to confer new biochemical properties to the IS110 system. In this study, IS110 transposases with unusually long amino acid sequences can be investigated as potential candidates for these domain fusions. In some embodiments, protein domain collections such as Pfam or InterPro can be used to search for all proteins known to contain the IS110 RuvC-like DEDD catalytic domain, the IS110 transposase domain, or both. These IS110 domains can then be synthesized and experimentally characterized to identify any favorable properties they may possess. In another embodiment, IS110 transposases may not be fused to additional protein cofactors as a single amino acid sequence, but may associate with them in multiprotein complexes. These complexes can be identified by searching for proteins that appear near the IS110 CDS at high frequency in natural genomes. These additional cofactors can be synthesized and experimentally characterized to identify any favorable properties they may possess.

[0152] In some embodiments, the present invention contemplates approaches to engineering the IS110 system in which the transposase domain is replaced or fused with an integrase, nucleobase deaminase, reverse transcriptase, recombinase, integrase, topoisomerase, retrotransposon, phosphatase, polymerase, ligase, helitron, helicase, methylase, demethylase, translation activator, translation repressor, transcription activator, transcription repressor, transcription dissociator, chromatin modifier, histone modifier, acetylase, deacetylase, reverse transcriptase, or nuclease. Fusion to a transposase domain can include simply fusing any of the aforementioned effector domain-containing enzymes or their domain(s) to either the N- or C-terminus of the transposase. In some embodiments, the IS110 transposase can also be modified such that the catalytic activity of RuvC or the transposase domain is inactivated. In some embodiments, an alternative to direct fusion is to co-deliver one or more of these effector domains, for example, along with an inactive IS110 transposase domain.

[0153] In some embodiments, the present invention contemplates an approach to engineering the IS110 system in which the IS110 transposase contains one or more amino acid mutations compared to the wild-type, whereby the mutations increase binding and / or interaction with the target site sequence, the donor site sequence, and / or the bridge RNA, and / or increase IS110 transposase activity.

[0154] In some embodiments, following the observation that IS110 transposases excise themselves using inefficient mechanisms, IS110 transposases can be engineered to reduce, limit, or eliminate insertion, excision recombination, or inversion reversal reactions.

[0155] In certain aspects, described herein are nucleic acids encoding the amino acid sequence of any of the IS110 family transposases provided herein. In any of the embodiments described herein, the nucleotide sequence encoding the IS110 transposase or a portion thereof can be codon-optimized. This type of optimization is known in the art and involves mutating exogenous DNA to mimic the codon preferences of the intended host organism or cell while still encoding the same protein. Thus, the codons are altered, but the encoded protein is not. For example, if the intended target cell is a human cell, a human codon-optimized IS110 transposase or a portion thereof is a suitable IS110 transposase. As another non-limiting example, if the intended host cell is a mouse cell, a mouse codon-optimized IS110 transposase or a portion thereof is a suitable IS110 transposase. Codon optimization is not required, but is acceptable and may be preferred in certain cases.

[0156] D. Bridge RNA It has been bioinformatically and experimentally determined that the IS110 transposase binds to an RNA, referred to herein as the bridge RNA, which is expressed from the non-coding end of the IS110 element and directs donor and target site specificity. See Examples 1 and 3-5. The bridge RNA was previously referred to as the "directing RNA" or "dRNA," and all references to the "directing RNA" or "dRNA" refer to the bridge RNA.

[0157] As discussed above, the linkage of the RE and LE can form a promoter that drives expression of the bridge RNA. The bridge RNA forms an RNA-protein complex with the transposase, recognizing the target site or the donor and target sites together through base pairing and mediating transposition. In some embodiments, the bridge RNA can be recombinantly expressed using any suitable promoter. In some embodiments, portions of the bridge RNA can be recombinantly expressed as two separate molecules using a single promoter or multiple promoters. In some embodiments, the bridge RNA can be recombinantly expressed using a promoter designed from the linked sequence of the RE and LE as described above. In some embodiments, the bridge RNA can be recombinantly expressed using a type III pol III promoter known in the art to be suitable for expressing small RNAs, such as guide RNAs, in CRISPR-Cas systems. In some embodiments, the promoter is a sigma 70 promoter box.

[0158] Bridge RNAs for IS110 transposases of the IS110 group are typically encoded within the LE, but in some embodiments, the bridge RNA can also be at least partially encoded within the RE. In some embodiments, the bridge RNA can be encoded within the LE and at least partially encoded within the 5' portion of the CDS sequence encoding the IS110 transposase. Bridge RNAs for IS110 transposases of the IS110 group have donor and target binding loops that recognize donor and target site sequences through base pairing to mediate transposition.

[0159] Bridge RNAs for IS1111 group IS110 transposases are typically encoded within an RE. In some embodiments, the bridge RNA can be encoded within an RE and at least partially within the 3' portion of a CDS sequence encoding an IS110 transposase (e.g., an IS1111 group IS110 transposase). Bridge RNAs for IS1111 group IS110 transposases have donor and target binding loops that recognize the donor and target site sequences through base pairing to mediate transposition. In some embodiments, the donor and target binding sequences of a bridge RNA for IS1111 group IS110 transposases can be present on a single loop of the bridge RNA that recognizes the donor and target site sequences through base pairing to mediate transposition. In some embodiments, the donor binding sequence of a bridge RNA for IS1111 group IS110 transposases can be present on a multi-branched loop of the bridge RNA that recognizes the donor site sequence through base pairing to mediate transposition.

[0160] Bridge RNAs include stem structures with one or more stem-loop structures (external loops, bulge loops, multi-branched loops, hairpin loops, and / or internal loops) formed through intramolecular base pairing. See, e.g., Sato, K., Akiyama, M. & Sakakibara, Y. RNA secondary structure prediction using deep learning with thermodynamic integration. Nat Commun 12, 941 (2021), see Figure 1 in doi.org / 10.1038 / s41467-021-21194-4 " refers to a single-stranded RNA molecule that forms a stem (or a nucleotide sequence) of a nucleic acid molecule, the contents of which are incorporated herein by reference in their entirety. Typically, the stem is formed by two portions of the RNA molecule that are complementary and base-pair to form the stem when read in reverse, with unpaired nucleotides intervening between the two complementary portions forming a loop. The stem sequence need not be 100% complementary and may contain mismatches, bulges, or internal loops.

[0161] In some embodiments, the bridge RNA comprises an RNA molecule comprising at least one stem-loop structure and further comprising one or more loops, the loops comprising: a first nucleotide sequence complementary to a first target site sequence of the target DNA; a second nucleotide sequence complementary to a second target site sequence on the opposite strand of the target DNA from the first target site sequence; a third nucleotide sequence complementary to a first donor site sequence of the donor DNA; and a fourth nucleotide sequence complementary to a second donor site sequence on the opposite strand of the donor DNA from the first donor site sequence. In some embodiments, the bridge RNA comprises a first internal loop comprising a first nucleotide sequence complementary to a first target site sequence of the target DNA and a second nucleotide sequence complementary to a second target site sequence on the opposite strand of the target DNA from the first target site sequence; and a second internal loop comprising a third nucleotide sequence complementary to a first donor site sequence of the donor DNA and a fourth nucleotide sequence complementary to a second donor site sequence on the opposite strand of the donor DNA from the first donor site sequence. In some embodiments, the bridge RNA comprises an internal loop comprising a first nucleotide sequence complementary to a first target site sequence of the target DNA and a second nucleotide sequence complementary to a second target site sequence on the opposite strand of the target DNA from the first target site sequence, and a multi-branched loop comprising a third nucleotide sequence complementary to a first donor site sequence of the donor DNA and a fourth nucleotide sequence complementary to a second donor site sequence on the opposite strand of the donor DNA from the first donor site sequence. In some embodiments, the bridge RNA binds to an IS110 group transposase. In some embodiments, the bridge RNA binds to an IS1111 group transposase. In some embodiments, the bridge RNA binds to IS621, ISPa11, IsPa29, ISMmg1, ISPfl1, ISMae40, ISStma6, ISAzs32, ISMex9, ISCARP28, ISAar16, ISCps7, ISPpu9, ISRel9, ISEsa2, ISMma5, IS900, or ISHne5 transposase.In some embodiments, the bridge RNA binds to IS621 transposase. In some embodiments, a loop comprising first and second nucleotide sequences complementary to a target site sequence of a target DNA can form a stem or partial stem structure when not bound to a transposase and a single-stranded structure when bound to a transposase. In some embodiments, a loop comprising third and fourth nucleotide sequences complementary to a donor site sequence of a donor DNA can form a stem or partial stem structure when not bound to a transposase and a single-stranded structure when bound to a transposase.

[0162] With respect to the bridge RNA described above, in some embodiments, the first and second nucleotide sequences are fully complementary to their respective target site sequences in the target DNA. In some embodiments, the first and / or second nucleotide sequences are partially complementary to their respective target site sequences in the target DNA, i.e., there is tolerance for non-canonical base pairing, mismatches, and / or non-contiguous sequences. In some embodiments, there is a single mismatch in the first or second nucleotide sequence. In some embodiments, there are two, three, or four mismatches, which may be within the first nucleotide sequence, within the second nucleotide sequence, or may be spread across the first and second nucleotide sequences. In some embodiments, the two, three, or four mismatches are contiguous. In some embodiments, there is a single non-canonical base pair in the first or second nucleotide sequence. In some embodiments, there is a single non-canonical base pair in the first nucleotide sequence, within the second nucleotide sequence, or spread across the first and second nucleotide sequences. There are two, three, or four non-canonical base pairs that may be within the first and second nucleotide sequences, or that may be spread across the first and second nucleotide sequences. In some embodiments, the two, three, or four non-canonical base pairs are contiguous. In some embodiments, the two, three, or four non-canonical base pairs are non-contiguous. In some embodiments, the third and fourth nucleotide sequences are fully complementary to the respective donor site sequences of the donor DNA. In some embodiments, the third and / or fourth nucleotide sequences are partially complementary to the respective donor site sequences of the donor DNA, i.e., there is tolerance for non-canonical base pairing, mismatches, and / or non-contiguous sequences. In some embodiments, there is a single mismatch within the third or fourth nucleotide sequence. In some embodiments, there are two mismatches that may be within the third nucleotide sequence, the fourth nucleotide sequence, or that may be spread across the third and fourth nucleotide sequences. In some embodiments, the two mismatches are contiguous. In some embodiments, the two mismatches are non-contiguous. In some embodiments, there is a single non-canonical base pair within the third or fourth nucleotide sequence. In some embodiments, there are two non-canonical base pairs, which may be within the third nucleotide sequence, the fourth nucleotide sequence, or may span the third and fourth nucleotide sequences. In some embodiments, the two non-canonical base pairs are contiguous. In some embodiments, the two non-canonical base pairs are non-contiguous. In some embodiments, the non-canonical base pairing is non-Watson-Crick base pairing (i.e., not a GC base pair or an AT / U base pair). See, e.g., Pacesa M, Lin CH, Clery A, et al., Structural basis for Cas9 off-target activity, Cell, 2022, 185(22):4067-4081.e21 (the contents of which are incorporated herein by reference in their entirety). In some embodiments, the non-canonical base pairing is a wobble base pairing. In some embodiments, the non-canonical base pairing is Hoogsteen base pairing.In some embodiments, the non-canonical base pairing comprises an rG-dT base pair, an rU-dG base pair, an rA-dC base pair, an rC-dA base pair, an rA-dG base pair, or an rG-dG base pair.

[0163] In some embodiments, the bridge RNA comprises an RNA molecule comprising at least one stem-loop structure and further comprises one or more loops comprising a first nucleotide sequence complementary to a first target site sequence of a target DNA and a second nucleotide sequence complementary to a second target site sequence on the opposite strand of the target DNA from the first target site sequence. In some embodiments, the bridge RNA further comprises a third nucleotide sequence complementary to a first donor site sequence of a donor DNA and a fourth nucleotide sequence complementary to a second donor site sequence on the opposite strand of the donor DNA from the first donor site sequence. In some embodiments, the bridge RNA comprises a first internal loop comprising a first nucleotide sequence complementary to a first target site sequence of a target DNA and a second nucleotide sequence complementary to a second target site sequence on the opposite strand of the target DNA from the first target site sequence. In some embodiments, the bridge RNA comprises a second internal loop comprising a third nucleotide sequence complementary to a first donor site sequence of donor DNA and a fourth nucleotide sequence complementary to a second donor site sequence on the opposite strand of the donor DNA from the first donor site sequence. In some embodiments, the first internal loop (e.g., a multi-branched loop) of the bridge RNA comprises a third nucleotide sequence complementary to a first donor site sequence of donor DNA and a fourth nucleotide sequence complementary to a second donor site sequence on the opposite strand of the donor DNA from the first donor site sequence. In some embodiments, the bridge RNA binds to an IS1111 group transposase. In some embodiments, the bridge RNA binds to an IS1111_229727 transposase. In some embodiments, the loop comprising the first and second nucleotide sequences complementary to target site sequences of target DNA forms a stem or partial stem structure when not bound to a transposase, and is capable of translocating a target site sequence. In some embodiments, the loop comprising the third and fourth nucleotide sequences complementary to the donor site sequence of the donor DNA can form a stem or partial stem structure when not bound to the transposase and can form a single-stranded structure when bound to the transposase.

[0164] With respect to the bridge RNA described above, in some embodiments, the first and second nucleotide sequences are fully complementary to their respective target site sequences in the target DNA. In some embodiments, the first and / or second nucleotide sequences are partially complementary to their respective target site sequences in the target DNA, i.e., there is tolerance for non-canonical base pairing, mismatches, and / or non-contiguous sequences. In some embodiments, there is a single mismatch in the first or second nucleotide sequence. In some embodiments, there are two, three, or four mismatches, which may be within the first nucleotide sequence, the second nucleotide sequence, or may span the first and second nucleotide sequences. In some embodiments, the two, three, or four mismatches are contiguous. In some embodiments, the two, three, or four mismatches are non-contiguous. In some embodiments, there is a single non-canonical base pair in the first or second nucleotide sequence. In some embodiments, there are two, three, or four non-canonical base pairs, which may be within the first nucleotide sequence, the second nucleotide sequence, or may span the first and second nucleotide sequences. In some embodiments, the two, three, or four non-canonical base pairs are contiguous. In some embodiments, the two, three, or four non-canonical base pairs are non-contiguous. In some embodiments, the third and fourth nucleotide sequences are fully complementary to the respective donor site sequences of the donor DNA. In some embodiments, the third and / or fourth nucleotide sequences are partially complementary to the respective donor site sequences of the donor DNA, i.e., there is tolerance for non-canonical base pairing, mismatches, and / or non-contiguous sequences. In some embodiments, there is a single mismatch in the third or fourth nucleotide sequence. In some embodiments, there are two mismatches, which may be within the third nucleotide sequence, the fourth nucleotide sequence, or may span the third and fourth nucleotide sequences. In some embodiments, the two mismatches are contiguous. In some embodiments, the two mismatches are non-consecutive.In some embodiments, there is a single non-canonical base pair in the third or fourth nucleotide sequence. In some embodiments, there are two non-canonical base pairs, which may be within the third nucleotide sequence, the fourth nucleotide sequence, or may span the third and fourth nucleotide sequences. In some embodiments, the two non-canonical base pairs are contiguous. In some embodiments, the two non-canonical base pairs are non-contiguous. In some embodiments, the non-canonical base pairing is non-Watson-Crick base pairing (i.e., not a GC base pair or an AT / U base pair). In some embodiments, the non-canonical base pairing is a wobble base pairing. In some embodiments, the non-canonical base pairing is a Hoogsteen base pairing. In some embodiments, the non-canonical base pairing comprises an rG-dT base pair, an rU-dG base pair, an rA-dC base pair, an rC-dA base pair, an rA-dG base pair, or an rG-dG base pair.

[0165] In some embodiments, the bridge RNA comprises an RNA molecule comprising at least two stem-loop structures; for a bridge RNA having at least two stem-loop structures, these two structures are referred to as a "first stem-loop" and a "second stem-loop," where the first is 5' or upstream of the second. In some embodiments, the first stem-loop comprises a target-binding loop and the second stem-loop comprises a donor-binding loop. In some embodiments, the first stem-loop comprises a donor-binding loop and the second stem-loop comprises a target-binding loop. Thus, in this nomenclature, additional stem-loop structures may be present upstream of the first stem-loop, between the first and second stem-loops, and / or after the second stem-loop. In some embodiments, the first stem-loop comprises a target-binding loop and the second stem-loop comprises a donor-binding loop. The bridge RNA comprises at least a first internal loop, termed the target-binding loop, and a second internal loop, termed the donor-binding loop. In some embodiments, the bridge RNA comprises an RNA molecule comprising at least two stem-loop structures, as illustrated in FIG. 2D and for cluster 1 of FIG. 13. In FIG. 2D and for cluster 1 of FIG. 13, and in some embodiments, the first stem-loop structure of the bridge RNA comprises a first stem-loop (5-35 nt, 3-10 nt loop) comprising an internal loop (e.g., target-binding loop) (5-20 nt). In FIG. 2D and for cluster 1 of FIG. 13, and in some embodiments, the second stem-loop structure of the bridge RNA comprises a second stem-loop (5-35 nt, 3-10 nt loop) comprising an internal loop (e.g., donor-binding loop) (5-20 nt). The stem of the second stem-loop structure may comprise an additional loop and bubble, each of 1-10 nucleotides. As shown in FIG. 2D , and in some embodiments, the bridge RNA may further comprise an additional third stem-loop structure (5-15 nt stem, 3-10 nt loop) (accessory structure) 5′ to the first stem-loop. In some embodiments, the bridge RNA binds to IS621 transposase. In some embodiments, the bridge RNA binds to ISMae40 or ISMetma6 transposase. In some embodiments, the donor binding loop and / or the target binding loop may form a stem or partial stem structure when not bound to a transposase, and a single-stranded structure when bound to a transposase.

[0166] Thus, in some embodiments, the bridge RNA comprises a nucleotide sequence comprising the following secondary structure: a first stem-loop comprising an internal target binding loop—a second stem-loop comprising an internal donor binding loop, hi some embodiments, the bridge RNA comprises additional stem-loop structures, bulges, and / or loops (see, e.g., Figure 2D). In some embodiments, the bridge RNA comprises a nucleotide sequence comprising: 5'-[A]-[B]-[C]-[D]-[E]-[F]-[G]-[H]-[I]-[J]-[K]-[L]-[M]-[N]-3', where A is a first stem portion, B is a first side of the internal loop corresponding to the target binding loop, C is a second stem portion, D is a first loop portion, E is the reverse complement of C, F is a second side of the internal loop corresponding to the target binding loop, G is the reverse complement of A, H is a third stem portion, I is a first side of the internal loop corresponding to the donor binding loop, J is a fourth stem portion, K is a second loop portion, L is the reverse complement of J, M is a second side of the internal loop corresponding to the donor binding loop, and N is the reverse complement of H. In some embodiments, the reverse complement sequence portions are not 100% complementary, such that the stem structure may contain one or more mismatches or bulges, or non-canonical base pairing may occur. In some embodiments, a first side of the internal loop corresponding to the target binding loop comprises a nucleotide sequence complementary to a first target site sequence of the target DNA (referred to as LTG (left target guide)), and a second side of the internal loop corresponding to the target binding loop comprises a second nucleotide sequence complementary to a second target site sequence on the opposite strand of the target DNA from the first target site sequence (referred to as RTG (right target guide)). In some embodiments, a first side of the internal loop corresponding to the donor binding loop comprises a nucleotide sequence complementary to a first donor site sequence of the donor DNA (referred to as LDG (left donor guide)), and a second side of the internal loop corresponding to the donor binding loop comprises a second nucleotide sequence complementary to a second donor site sequence on the strand of the donor DNA opposite the first target site sequence (referred to as RDG (right donor guide)). In some embodiments, the stem structure may comprise one or more mismatches or bulges.In some embodiments, at least two portions of nucleotide sequence N are not complementary to the nucleotide sequence of portion H, such that the stem structure formed by base pairing between portions H and N includes two bulges. In some embodiments, additional nucleotides are present between any portion of A through N. For example, one or more non-base-pairing nucleotides are present between portions G and H. In other embodiments, one or more nucleotides are present between the 5' end and portion A. In some embodiments, one or more nucleotides are present between the 3' end and the portion N. In some embodiments, the bridge RNA binds to an IS621 transposase. In some embodiments, the bridge RNA binds to an ISMae40 or an ISMetma6 transposase. In some embodiments, the donor binding loop and / or the target binding loop may form a stem or partial stem structure when not bound to a transposase, and a single-stranded structure when bound to a transposase.

[0167] In some embodiments, the bridge RNA comprises a nucleotide sequence comprising the following secondary structure: third stem loop—first stem loop comprising an internal target binding loop—second stem loop comprising an internal donor binding loop. In some embodiments, the bridge RNA comprises an additional stem loop structure, bulge, and / or loop (see, e.g., Figure 2D). In some embodiments, the bridge RNA comprises a nucleotide sequence comprising 5'-[Z]-[X]-[Y]-[A]-[B]-[C]-[D]-[E]-[F]-[G]-[H]-[I]-[J]-[K]-[L]-[M]-[N]-3', where Z is the fifth stem portion, X is the third loop portion, Y is the reverse complement of Z, A is the first stem portion, and B is the first internal loop portion corresponding to the target binding loop. side of the internal loop corresponding to the target binding loop, C is a second stem portion, D is a first loop portion, E is the reverse complement of C, F is a second side of the internal loop corresponding to the target binding loop, G is the reverse complement of A, H is a third stem portion, I is a first side of the internal loop corresponding to the donor binding loop, J is a fourth stem portion, K is a second loop portion, L is the reverse complement of J, M is a second side of the internal loop corresponding to the donor binding loop, and N is the reverse complement of H. In some embodiments, the reverse complement portions are not 100% complementary, so the stem structure may contain one or more mismatches or bulges, or non-canonical base pairing may occur. In some embodiments, a first side of the internal loop corresponding to the target binding loop comprises a nucleotide sequence (referred to as LTG) complementary to a first target site sequence of the target DNA, and a second side of the internal loop corresponding to the target binding loop comprises a second nucleotide sequence (referred to as RTG) complementary to a second target site sequence on the opposite strand of the target DNA from the first target site sequence. In some embodiments, a first side of the internal loop corresponding to the donor binding loop comprises a nucleotide sequence (referred to as LDG) complementary to a first donor site sequence of the donor DNA, and a second side of the internal loop corresponding to the donor binding loop comprises a second nucleotide sequence (referred to as RDG) complementary to a second donor site sequence on the opposite strand of the donor DNA from the first target site sequence. In some embodiments, the stem structure may comprise one or more mismatches or bulges.In some embodiments, at least two portions of nucleotide sequence N are not complementary to the nucleotide sequence of portion H, such that the stem structure formed by base pairing between portions H and N includes two bulges. In some embodiments, additional nucleotides are present between any portion A through N. For example, one or more non-base-pairing nucleotides are present between portions G and H. In another embodiment, one or more nucleotides are present between portions Y and A. In another embodiment, one or more nucleotides are present between the 5' end and portion Z. In another embodiment, one or more nucleotides are present between the 3' end and portion N. In some embodiments, the bridge RNA binds to IS621 transposase. In some embodiments, the donor binding loop and / or the target binding loop may form a stem or partial stem structure when not bound to the transposase, and a single-stranded structure when bound to the transposase.

[0168] With respect to the bridge RNAs described above, in some embodiments, the LTG and RTG sequences are fully complementary to their respective target site sequences in the target DNA. In some embodiments, the LTG and / or RTG are partially complementary to their respective target site sequences in the target DNA, i.e., there is tolerance for non-canonical base pairing, mismatches, and / or non-consecutive base pairing. In some embodiments, there is a single mismatch in the LTG or RTG nucleotide sequence. In some embodiments, there are two, three, or four mismatches. In some embodiments, the two, three, or four mismatches are contiguous. In some embodiments, the two, three, or four mismatches are non-contiguous. In some embodiments, the LTG and / or RTG are partially complementary to their respective target site sequences in the target DNA, such that one, two, three, or four nucleotides dispersed within the LTG and / or RTG do not base pair with the target DNA, forming a bulge. In some embodiments, there is a single non-canonical base pair in the LTG or RTG nucleotide sequence. In some embodiments, there are two, three, or four non-canonical base pairs, which may be within the LTG, within the RTG, or spread across the LTG and RTG. In some embodiments, the two, three, or four non-canonical base pairs are contiguous. In some embodiments, the two, three, or four non-canonical base pairs are non-contiguous. In some embodiments, the LDG and RDG nucleotide sequences are fully complementary to the respective donor site sequences of the donor DNA. In some embodiments, the LDG and / or RDG nucleotide sequences are partially complementary to the respective donor site sequences of the donor DNA, i.e., there is tolerance for non-canonical base pairing, mismatches, and / or non-contiguous sequences. In some embodiments, there is a single mismatch in the LDG or RDG nucleotide sequence. In some embodiments, there are two, three, or four mismatches, which may be within the LDG, within the RDG, or may span the LDG and RDG. In some embodiments, the two, three, or four mismatches are contiguous. In some embodiments, the two, three, or four mismatches are non-contiguous. In some embodiments, the LDG and / or RDG are partially complementary to the respective donor site sequence of the donor DNA, such that 1, 2, 3, or 4 nucleotides dispersed within the LDG and / or RDG do not base pair with the donor DNA, forming a bulge (see, e.g., Figure 36A). In some embodiments, there is a single non-canonical base pair in the LDG or RDG nucleotide sequence.In some embodiments, there are two, three, or four non-canonical base pairs, which may be within the LDG, within the RDG, or spread across the LDG and the RDG. In some embodiments, the two, three, or four non-canonical base pairs are contiguous. In some embodiments, the two, three, or four non-canonical base pairs are non-contiguous. In some embodiments, the non-canonical base pairing is non-Watson-Crick base pairing (i.e., not a GC base pair or an AT / U base pair). In some embodiments, the non-canonical base pairing is a wobble base pairing. In some embodiments, the non-canonical base pairing is a Hoogsteen base pairing. In some embodiments, the non-canonical base pairing comprises an rG-dT base pair, an rU-dG base pair, an rA-dC base pair, an rC-dA base pair, an rA-dG base pair, or an rG-dG base pair.

[0169] In some embodiments, the bridge RNA comprises a nucleotide sequence comprising the sequence 5'--nnnnnnnnYYnRRnn----nnYYYYnnnYnnnnRRnnnnYYGGAYGCCGYnYYnRnCCUnnRRYnnnARYYYGYnnYGUAGAUnnnYGCRnC-RRnYRYYnnnnnnnnYnnGYnnRRRYCGRACnGnAUCnYnGGCYGGY-nnnYCGRnARYCYGCAUUYACAAGGUnGRUnRCRYRAnnn-n 3' (SEQ ID NO: 795156), where "n" represents any nucleotide, "R" represents an A or G nucleotide, and "Y" represents a C or U nucleotide. In some embodiments, the bridge RNA comprises a nucleotide sequence comprising the sequence 5'--nnnnnnnnYYnRRnn----nnYYYnnnYnnnnRRnnnnYYGGAYGCCGYnYYnRnCCUnnRRYnnnARYYYGYnnYGUAGAUnnnYGCRnC-RRnYRYYnnnnnnnnYnnGYnnRRRYCGRACnGnAUCnYnGGCYGGY-nnnYCGRnARYCYGCAUUYACAAGGUnGRUnRCRYRAnnn-n 3' (SEQ ID NO: 795156), where "n" represents any nucleotide. wherein "R" represents an A or G nucleotide and "Y" represents a C or U nucleotide, and wherein the bridge RNA comprises the secondary structure 5'........((((((..........)))))).....(((((((((.(((............(((((....)))))..............))))))))))...................................(((((.....)))))..................................3', where matching parentheses "(" and ")" indicate base-paired nucleotides and "." indicates an unpaired base. In some embodiments, the bridge RNA comprises a nucleotide sequence comprising the sequence 5'--nnnnnnnnYYnRRnn----nnYYYYnnnYnnnnRRnnnnYYGGAYGCCGYnYYnRnCCUnnRRYnnnARYYYGYnnYGUAGAUnnnYGCRnC-RRnYRYYnnnnnnnnYnnGYnnRRRYCGRACnGnAUCnYnGGCYGGY-nnnYCGRnARYCYGCAUUYACAAGGUnGRUnRCRYRAnnn-3' (SEQ ID NO: 795156), where "n" represents any nucleotide and "R" represents an A or G nucleotide. and "Y" represents a C or U nucleotide, wherein the bridge RNA comprises the secondary structure 5'.....((((((((((........)))))))))).((((((((((.(((............((((((....)))))..............))))))))))).......(.((((....(((.............(((((((.....))))))...)...............))))))))....3', where matching brackets "(" and ")" indicate base-paired nucleotides and "." indicates an unpaired base.In some embodiments, the bridge RNA comprises a nucleotide sequence comprising the sequence 5'--nnnnnnnnYYnRRnn----nnYYYYnnnYnnnnRRnnnnYYGGAYGCCGYnYYnRnCCUnnRRYnnnARYYYGYnnYGUAGAUnnnYGCRnC-RRnYRYYnnnnnnnnYnnGYnnRRRYCGRACnGnAUCnYnGGCYGGY-nnnYCGRnARYCYGCAUUYACAAGGUnGRUnRCRYRAnnn-3' (SEQ ID NO: 795156), where "n" represents any nucleotide and "R" represents an A or G nucleotide. and "Y" represents a C or U nucleotide, wherein the bridge RNA comprises the secondary structure 5'.....(((((((((((......))))))))))).(((((((((((.((...(((.((...(((.(...((((((....)))))..))...))))....)))))))))))......((((((((((..((((...........((((((((.....)))))...)))).............))))))))))))...3', where matching brackets "(" and ")" indicate base-paired nucleotides and "." indicates an unpaired base.In some embodiments, the bridge RNA comprises a nucleotide sequence comprising the sequence 5'--nnnnnnnnYYnRRnn----nnYYYnnnYnnnnRRnnnnYYGGAYGCCGYnYYnRnCCUnnRRYnnnARYYYGYnnYGUAGAUnnnYGCRnC-RRnYRYYnnnnnnnnYnnGYnnRRRYCGRACnGnAUCnYnGGCYGGY-nnnYCGRnARYCYGCAUUYACAAGGUnGRUnRCRYRAnnn-n 3' (SEQ ID NO: 795156), where "n" represents any nucleotide and "R" represents an A or G nucleotide. and "Y" represents a C or U nucleotide, and wherein the bridge RNA comprises the secondary structure 5'.....((((((((((......)))))))))).((((((((((.((...((((((((((....)))))..))))..)))))))))...((.(((((((((..((((.((....((((((((....))))))..) ... Combine with 1.

[0170] In some embodiments, the bridge RNA comprises at least a target binding loop and a donor binding loop and is encoded by a sequence that is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the bridge RNA sequence of Figure 15 (SEQ ID NOs: 1-348) or SEQ ID NOs: 349-10175. In some embodiments, the bridge RNA comprises a sequence that is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the bridge RNA sequence of Figure 15 (SEQ ID NOs: 1-348) or SEQ ID NOs: 349-10175.

[0171] In some embodiments, the bridge RNA comprises a nucleotide sequence comprising any of the 5' to 3' sequences shown in Figure 19, where "n" represents any nucleotide, "R" represents an A or G nucleotide, and "Y" represents a C or U nucleotide. In some embodiments, the bridge RNA comprises a nucleotide sequence comprising any of the 5' to 3' sequences shown in Figure 19, where "n" represents any nucleotide, "R" represents an A or G nucleotide, and "Y" represents a C or U nucleotide, where the bridge RNA comprises a 5' to 3' secondary structure as shown in the first column of the secondary structure of said sequence shown in Figure 19, where matching brackets "(" and ")" indicate base-paired nucleotides and "." indicates an unpaired base. In some embodiments, the bridge RNA comprises a nucleotide sequence comprising any of the 5' to 3' sequences shown in Figure 19, where "n" represents any nucleotide, "R" represents an A or G nucleotide, and "Y" represents a C or U nucleotide, wherein the bridge RNA comprises the 5' to 3' secondary structure shown in the second column of the secondary structure of said sequence provided in Figure 19, where matching brackets "(" and ")" indicate base-paired nucleotides and "." indicates an unpaired base. In some embodiments, the bridge RNA comprises a nucleotide sequence comprising any of the 5' to 3' sequences shown in Figure 19, where "n" represents any nucleotide, "R" represents an A or G nucleotide, and "Y" represents a C or U nucleotide, wherein the bridge RNA comprises the 5' to 3' secondary structure shown in the third column of the secondary structure of said sequence provided in Figure 19, where matching brackets "(" and ")" indicate base-paired nucleotides and "." indicates an unpaired base. In some embodiments, the bridge RNA comprises a nucleotide sequence comprising any of the 5' to 3' sequences shown in Figure 19, where "n" represents any nucleotide, "R" represents an A or G nucleotide, and "Y" represents a C or U nucleotide, and wherein the bridge RNA comprises a 5' to 3' secondary structure shown in the fourth column of the secondary structure of said sequence shown in Figure 19, where matching parentheses "(" and ")" indicate base-paired nucleotides and "." indicates an unpaired base.In some embodiments, the bridge RNA binds to the IS110 transposase indicated after ">" for the sequence shown in Figure 19. In some embodiments, the nucleotide sequence or nucleotide sequence and secondary structure from Figure 19 is for the bridge RNA of ISPa11, IsPa29, ISMmg1, ISPfl1, ISMae40, ISStma6, ISAzs32, ISMex9, ISCARP28, ISAar16, ISCps7, ISPpu9, ISRel9, ISEsa2, ISMma5, IS900, or ISHne5.

[0172] The target binding loop of the first stem-loop structure comprises a nucleotide sequence designated as a left targeting guide (LTG) and a nucleotide sequence designated as a right targeting guide (RTG). In some embodiments, the LTG and RTG sequences can be reprogrammed to bind to a target site of interest. See Section E. In some embodiments, the LTG and RTG sequences are not reprogrammed, and transpososomes containing such bridge RNAs are Target the wild-type target site sequence and the same sequence found in other organisms, or other sites similar to that site (e.g., sequences similar or identical to the wild-type target site sequence of the transposase). See Section E.

[0173] The donor binding loop of the second stem-loop structure comprises a nucleotide sequence designated as the left donor guide (LDG) and a nucleotide sequence designated as the right donor guide (RDG). In some embodiments, the LDG and RDG sequences can be reprogrammed to bind to a donor site sequence of interest. See Section E. In some embodiments, the LDG and RDG sequences are not reprogrammed, and transpososomes containing such bridge RNAs target the wild-type donor sequence and the same sequence found in other organisms, or other sites similar to that site (e.g., a sequence similar or identical to the wild-type donor site sequence of the transposase). See Section E.

[0174] Thus, the donor and target site specificity of IS110 can be encoded by the nucleotide sequences of the bridge RNA found within the target-binding loop and donor-binding loop. The sequence of at least one of LTG, RTG, LDG, or RDG can be reprogrammed via substitution, insertion, deletion, truncation, and extension (relative to the wild-type IS110 bridge RNA sequence). See Figure 4 and section E. Advantageously, described herein is that reprogramming of at least one of LTG, RTG, LDG, or RDG can confer specificity for any given sequence. The bridge RNA sequence can be reprogrammed via substitution, insertion, deletion, truncation, and extension. As described above, certain non-canonical base pairings, mismatches, and / or non-contiguous tolerances within the target-binding loop and / or donor-binding loop can be tolerated.

[0175] In some embodiments, the bridge RNA comprises an RNA molecule comprising a stem-loop structure comprising at least one multi-branched loop comprising at least three stems. In some embodiments, the bridge RNA comprises at least a first internal loop, termed the target binding loop, comprising RTG and LTG sequences, and a second internal loop, termed the donor binding loop, comprising RDG and LDG sequences. In some embodiments, the bridge RNA comprises an RNA molecule comprising a stem-loop structure comprising at least one multi-branched loop comprising at least three stems, as shown in Cluster 1 of Figure 13. In Cluster 1 of Figure 13 and in some embodiments, the stem-loop structure of the bridge RNA comprises a multi-branched loop comprising a first stem, a second stem comprising a stem-loop and an internal loop, and a third stem comprising a stem-loop. The first stem, the second stem, or the third stem may comprise an additional loop and a bubble. One or more of the loops (e.g., a hyperbranched loop, or one or more internal loops, e.g., an internal loop of the second stem) may correspond to a target-binding loop and a donor-binding loop and may comprise nucleotide sequences corresponding to RTG and LTG sequences and RDG and LDG sequences. In some embodiments, the internal loop of the second stem comprises a target-binding loop, and the hyperbranched loop comprises a donor-binding loop. See, e.g., Figures 35A-B. As shown in cluster 1 of Figure 13, and in some embodiments, the bridge RNA may further comprise an additional stem-loop structure 5' to the hyperbranched loop structure. In some embodiments, the bridge RNA binds to ISPa11, ISPa29, ISMmg1, or ISPfl1 transposase. In some embodiments, the donor-binding loop and / or the target-binding loop may form a stem or partial stem structure when not bound to a transposase, and may form a single-stranded structure when bound to a transposase.

[0176] In some embodiments, the bridge RNA comprises an RNA molecule comprising a stem-loop structure comprising at least one multi-branched loop comprising at least three stems. In some embodiments, the bridge RNA comprises a small branched loop, referred to as a target binding loop, comprising RTG and LTG sequences. The bridge RNA comprises at least a first internal loop and a second internal loop, referred to as the donor binding loop, that comprises the RDG and LDG sequences. In some embodiments, the bridge RNA comprises an RNA molecule comprising a stem-loop structure comprising at least one multi-branched loop comprising at least three stems, as shown in cluster 3 of Figure 13. In cluster 3 of Figure 13 and in some embodiments, the stem-loop structure of the bridge RNA comprises a multi-branched loop comprising a first stem, a second stem comprising a stem-loop and an internal loop, and a third stem comprising a stem-loop. The first stem, the second stem, or the third stem may comprise additional loops and bubbles. One or more of the loops (e.g., the multi-branched loop, or one or more internal loops, e.g., the internal loop of the second stem) may correspond to the target binding loop and the donor binding loop and may comprise nucleotide sequences corresponding to the RTG and LTG sequences and the RDG and LDG sequences. In some embodiments, the internal loop of the second stem comprises the target binding loop, and the multi-branched loop comprises the donor binding loop. See, for example, Figures 35A-B. As shown in Cluster 3 of Figure 13, the bridge RNA further comprises an additional stem-loop structure 5' to the multi-branched loop structure. In some embodiments, the bridge RNA binds to an ISAzs32 or ISMex9 transposase. In some embodiments, the donor binding loop and / or the target binding loop can form a stem or partial stem structure when not bound to a transposase, and can form a single-stranded structure when bound to a transposase.

[0177] In some embodiments, the bridge RNA comprises an RNA molecule comprising a stem-loop structure comprising at least one multi-branched loop comprising at least three stems. In some embodiments, the bridge RNA comprises at least a first internal loop, termed the target binding loop, comprising RTG and LTG sequences, and a second internal loop, termed the donor binding loop, comprising RDG and LDG sequences. In some embodiments, the bridge RNA comprises an RNA molecule comprising a stem-loop structure comprising at least one multi-branched loop comprising at least three stems, as shown in cluster 4 of Figure 13. In cluster 4 of Figure 13 and in some embodiments, the stem-loop structure of the bridge RNA comprises a multi-branched loop comprising a first stem, a second stem comprising a stem-loop and an internal loop, and a third stem comprising a stem-loop and an internal loop. The first stem, the second stem, or the third stem may comprise additional loops and bubbles. One or more of the loops (e.g., a multi-branched loop, or one or more internal loops, e.g., an internal loop of the second stem or an internal loop of the third stem) may correspond to a target binding loop and a donor binding loop and may comprise nucleotide sequences corresponding to RTG and LTG sequences and RDG and LDG sequences. In some embodiments, the internal loop of the second stem comprises a target binding loop, and the internal loop of the third stem comprises a donor binding loop. See, e.g., Figures 35A-B. In some embodiments, the bridge RNA binds to an ISCARP28 transposase. In some embodiments, the donor binding loop and / or the target binding loop may form a stem or partial stem structure when not bound to a transposase, and may form a single-stranded structure when bound to a transposase.

[0178] In some embodiments, the bridge RNA comprises an RNA molecule comprising a stem-loop structure comprising a cloverleaf-like structure comprising at least three stem loops. In some embodiments, the bridge RNA comprises at least a first internal loop, termed the target binding loop, comprising RTG and LTG sequences, and a second internal loop, termed the donor binding loop, comprising RDG and LDG sequences. In some embodiments, the bridge RNA comprises an RNA molecule comprising a stem-loop structure comprising a cloverleaf-like structure comprising at least three stem loops, as shown in cluster 5 or cluster 12 of Figure 13. In cluster 5 or cluster 12 of Figure 13 and in some embodiments, the stem-loop structure of the bridge RNA comprises a cloverleaf-like structure comprising a first stem, a first stem loop, a second stem loop comprising an internal loop, and a third stem loop comprising an internal loop. The stem of the first, second, or third stem-loop may comprise an additional loop and a bubble. One or more of the loops (e.g., one or more internal loops, e.g., the internal loop of the second stem-loop or the internal loop of the third stem-loop) may correspond to a target binding loop and a donor binding loop and may comprise nucleotide sequences corresponding to RTG and LTG sequences and RDG and LDG sequences. In some embodiments, the internal loop of the second stem-loop comprises a target binding loop, and the internal loop of the third stem-loop comprises a donor binding loop. See, e.g., Figures 35A-B. In some embodiments, the bridge RNA binds to an ISAar16 or ISHne5 transposase. In some embodiments, the donor binding loop and / or the target binding loop may form a stem or partial stem structure when not bound to a transposase, and may form a single-stranded structure when bound to a transposase.

[0179] In some embodiments, the bridge RNA comprises an RNA molecule comprising a stem-loop structure comprising at least one multi-branched loop comprising at least three stems. In some embodiments, the bridge RNA comprises at least a first internal loop, termed the target binding loop, comprising RTG and LTG sequences, and a second internal loop, termed the donor binding loop, comprising RDG and LDG sequences. In some embodiments, the bridge RNA comprises an RNA molecule comprising a stem-loop structure comprising at least one multi-branched loop comprising at least three stems, as shown in Cluster 6 of Figure 13. In Cluster 6 of Figure 13 and in some embodiments, the stem-loop structure of the bridge RNA comprises a multi-branched loop comprising a first stem, a second stem comprising a stem-loop, and a third stem comprising a stem-loop and an internal loop. The first stem, the second stem, or the third stem may comprise additional loops and bubbles. One or more of the loops (e.g., a multi-branched loop, or one or more internal loops, e.g., an internal loop of the third stem) may correspond to a target binding loop and a donor binding loop and may comprise nucleotide sequences corresponding to RTG and LTG sequences and RDG and LDG sequences. In some embodiments, the internal loop of the third stem comprises a target binding loop, and the multi-branched loop comprises a donor binding loop. In some embodiments, the bridge RNA binds to a transposase of ISCps7. In some embodiments, the donor binding loop and / or the target binding loop may form a stem or partial stem structure when not bound to a transposase, and may form a single-stranded structure when bound to a transposase.

[0180] In some embodiments, the bridge RNA comprises an RNA molecule comprising a stem-loop structure comprising at least one multi-branched loop comprising at least three stems. In some embodiments, the bridge RNA comprises at least a first internal loop, termed the target binding loop, comprising RTG and LTG sequences, and a second internal loop, termed the donor binding loop, comprising RDG and LDG sequences. In some embodiments, the bridge RNA comprises an RNA molecule comprising a stem-loop structure comprising at least one multi-branched loop comprising at least three stems, as shown in Cluster 7 of Figure 13. In Cluster 7 of Figure 13 and in some embodiments, the stem-loop structure of the bridge RNA comprises a multi-branched loop comprising a first stem, a second stem comprising a stem-loop and an internal loop, and a third stem comprising a stem-loop and an internal loop. The first stem, the second stem, or the third stem may comprise additional loops and bubbles. One or more of the loops (e.g., a multi-branched loop, or one or more internal loops, e.g., an internal loop of the second stem or an internal loop of the third stem) may correspond to a target binding loop and a donor binding loop and may comprise nucleotide sequences corresponding to RTG and LTG sequences and RDG and LDG sequences. In some embodiments, the internal loop of the second stem comprises a target binding loop and the internal loop of the third stem comprises a donor binding loop. See, e.g., Figures 35A-B. In some embodiments, the bridge RNA binds to ISPpu9 or ISPpu10 transposase. In some embodiments, the donor binding loop and / or the target binding loop are It can form a stem or partial stem structure when not bound to a transposase, and can form a single-stranded structure when bound to a transposase.

[0181] In some embodiments, the bridge RNA comprises an RNA molecule comprising a stem-loop structure comprising at least one hyperbranched loop comprising at least four stems and at least two stem-loop structures. In some embodiments, the bridge RNA comprises at least a first internal loop, termed the target binding loop, comprising RTG and LTG sequences, and a second internal loop, termed the donor binding loop, comprising RDG and LDG sequences. In some embodiments, the bridge RNA comprises an RNA molecule comprising a stem-loop structure comprising at least one hyperbranched loop comprising at least four stems and at least two stem-loop structures, as shown in cluster 8 of Figure 13. In cluster 8 of Figure 13 and in some embodiments, the stem-loop structure of the bridge RNA comprises a first stem-loop, a second stem-loop comprising an internal loop, and a hyperbranched loop comprising a first stem, a second stem comprising a stem-loop, a third stem comprising a stem-loop, and a fourth stem comprising a stem-loop. The stems of the first stem-loop, the second stem-loop, and the first, second, third, or fourth stems of the hyperbranched loop may include additional loops and bubbles. One or more of the loops (e.g., the hyperbranched loop, or one or more internal loops, e.g., the internal loop of the second stem-loop) may correspond to the target-binding loop and the donor-binding loop and may include nucleotide sequences corresponding to the RTG and LTG sequences and the RDG and LDG sequences. In some embodiments, the internal loop of the second stem includes the target-binding loop, and the hyperbranched loop includes the donor-binding loop. In some embodiments, the bridge RNA binds to the transposase of ISRel9. In some embodiments, the donor-binding loop and / or the target-binding loop may form a stem or partial stem structure when not bound to the transposase, and may form a single-stranded structure when bound to the transposase.

[0182] In some embodiments, the bridge RNA comprises an RNA molecule comprising at least three stem-loop structures. In some embodiments, the bridge RNA comprises at least a first internal loop, termed the target binding loop, which comprises the RTG and LTG sequences, and a second internal loop, termed the donor binding loop, which comprises the RDG and LDG sequences. In some embodiments, the bridge RNA comprises an RNA molecule comprising a stem-loop structure comprising at least three stem-loop structures, as shown in cluster 9 or cluster 11 of Figure 13. In cluster 9 or cluster 11 of Figure 13 and in some embodiments, the bridge RNA comprises a first stem-loop, a second stem-loop comprising an internal loop, and a third stem-loop comprising an internal loop. The first stem-loop, the second stem-loop, or the third stem-loop may comprise an additional loop and a bubble. One or more of the loops (e.g., the internal loop of the second stem loop and / or the internal loop of the third stem loop) may correspond to a target binding loop and a donor binding loop and may comprise nucleotide sequences corresponding to RTG and LTG sequences and RDG and LDG sequences. In some embodiments, the internal loop of the second stem loop comprises a target binding loop, and the internal loop of the third stem loop comprises a donor binding loop. In some embodiments, the bridge RNA binds to an ISEsa2 or IS900 transposase. In some embodiments, the donor binding loop and / or the target binding loop may form a stem or partial stem structure when not bound to a transposase, and may form a single-stranded structure when bound to a transposase.

[0183] In some embodiments, the bridge RNA comprises an RNA molecule comprising a stem-loop structure comprising at least one multi-branched loop comprising at least three stems. In some embodiments, the bridge RNA comprises at least a first internal loop, termed the target binding loop, comprising RTG and LTG sequences, and a donor binding loop, comprising RDG and LDG sequences. and a second internal loop. In some embodiments, the bridge RNA comprises an RNA molecule comprising a stem-loop structure comprising at least one multi-branched loop comprising at least three stems, as shown in cluster 10 of FIG. 13. In cluster 10 of FIG. 13 and in some embodiments, the stem-loop structure of the bridge RNA comprises a multi-branched loop comprising a first stem, a second stem comprising a stem-loop and an internal loop, and a third stem comprising a stem-loop and an internal loop. The first stem, the second stem, or the third stem may comprise additional loops and bubbles. One or more of the loops (e.g., a multi-branched loop, or one or more internal loops, e.g., an internal loop of the second stem or an internal loop of the third stem) may correspond to the target-binding loop and the donor-binding loop and may comprise nucleotide sequences corresponding to the RTG and LTG sequences and the RDG and LDG sequences. In some embodiments, the internal loop of the second stem comprises the target-binding loop, and the internal loop of the third stem comprises the donor-binding loop. In some embodiments, the internal loop of the second stem comprises a target binding loop and the multi-branched loop comprises a donor binding loop. In some embodiments, the bridge RNA binds to an ISMma5 transposase. In some embodiments, the donor binding loop and / or the target binding loop can form a stem or partial stem structure when not bound to the transposase, and can form a single-stranded structure when bound to the transposase.

[0184] With respect to any of the above bridge RNAs, in some embodiments, the LTG and RTG sequences are fully complementary to their respective target site sequences in the target DNA. In some embodiments, the LTG and / or RTG sequences are partially complementary to their respective target site sequences in the target DNA, i.e., there is tolerance for non-canonical base pairing, mismatches, and / or non-contiguous sequences. In some embodiments, there is a single mismatch in the LTG or RTG nucleotide sequence. In some embodiments, there are two, three, or four mismatches, which may be within LTG, within RTG, or may span LTG and RTG. In some embodiments, the two, three, or four mismatches are contiguous. In some embodiments, there is a single non-canonical base pair in the LTG or RTG nucleotide sequence. In some embodiments, there are two, three, or four non-canonical base pairs, which may be within LTG, within RTG, or may span LTG and RTG. In some embodiments, the two, three, or four non-canonical base pairs are contiguous. In some embodiments, the two, three, or four non-canonical base pairs are non-contiguous. In some embodiments, the LDG and RDG nucleotide sequences are fully complementary to the respective donor site sequences of the donor DNA. In some embodiments, the LDG and / or RDG nucleotide sequences are partially complementary to the respective donor site sequences of the donor DNA, i.e., there is tolerance for non-canonical base pairing, mismatches, and / or non-contiguous sequences. In some embodiments, there is a single mismatch in the LDG or RDG nucleotide sequence. In some embodiments, there are two, three, or four mismatches, which may be within the LDG, within the RDG, or may span the LDG and RDG. In some embodiments, the two, three, or four mismatches are contiguous. In some embodiments, the two, three, or four mismatches are non-contiguous. In some embodiments, there is a single non-canonical base pair in the LDG or RDG nucleotide sequence.In some embodiments, there are two, three, or four non-canonical base pairs, which may be within the LDG, the RDG, or may span the LDG and the RDG. In some embodiments, the two, three, or four non-canonical base pairs are contiguous. In some embodiments, the two, three, or four non-canonical base pairs are non-contiguous. In some embodiments, the non-canonical base pairing is non-Watson-Crick base pairing (i.e., not a GC base pair or an AT / U base pair). In some embodiments, the non-canonical base pairing is wobble base pairing. In some embodiments, the non-canonical base pairing is Hoogsteen base pairing. In some embodiments, the non-canonical base pairing is rG-dT, rU-dG, or rA-dC. base pair, rC-dA base pair, rA-dG base pair, or rG-dG base pair.

[0185] In some embodiments, the bridge RNA comprises an RNA molecule comprising a stem-loop structure comprising at least one multi-branched loop comprising at least three stems. In some embodiments, the bridge RNA comprises at least a first internal loop, termed the target binding loop, comprising RTG and LTG sequences, and a second internal loop, termed the donor binding loop, comprising RDG and LDG sequences. In some embodiments, the bridge RNA comprises an RNA molecule comprising a stem-loop structure comprising at least one multi-branched loop comprising at least three stems, as illustrated for IS1111 transposase in Figure 11B. In Figure 11B and in some embodiments, the stem-loop structure of the bridge RNA comprises a multi-branched loop comprising a first stem, a second stem comprising a stem-loop and an internal loop, and a third stem comprising a stem-loop and an internal loop. The first stem, second stem, or third stem may comprise additional loops and bubbles. One or more of the loops (e.g., the internal loop of the second stem and the internal loop of the third stem) may correspond to a target binding loop and a donor binding loop and may comprise nucleotide sequences corresponding to RTG and LTG sequences and RDG and LDG sequences. In some embodiments, the internal loop of the second stem comprises a target binding loop and the internal loop of the third stem comprises a donor binding loop. In some embodiments, the bridge RNA binds to the IS1111_229727 transposase. In some embodiments, the donor binding loop and / or the target binding loop may form a stem or partial stem structure when not bound to the transposase and may form a single-stranded structure when bound to the transposase.

[0186] Thus, in some embodiments, the bridge RNA comprises a nucleotide sequence comprising the following secondary structure: a first stem-multibranched loop-a second stem comprising a stem-loop and an internal loop-a third stem comprising a stem-loop and an internal loop. In some embodiments, the bridge RNA comprises additional stem-loop structures, bulges, and / or loops (see, e.g., Figure 11B). In some embodiments, the bridge RNA comprises a nucleotide sequence comprising: 5'-[A]-[B]-[C]-[D]-[E]-[F]-[G]-[H]-[I]-[J]-[K]-[L]-[M]-[N]-[O]-[P]-[Q]-[R]-[S]-3', where A is a first stem portion, B is a first portion of a multi-branched loop, C is a second stem portion, D is a first side of an internal loop corresponding to the target binding loop, E is a third stem portion, F is a first loop portion, and G is a second stem portion. where H is the second side of the internal loop corresponding to the target binding loop, I is the reverse complement of C, J is the second portion of the multi-branched loop, K is the fourth stem portion, L is the first side of the internal loop corresponding to the donor binding loop, M is the fifth stem portion, N is the second loop portion, O is the reverse complement of M, P is the second side of the internal loop corresponding to the donor binding loop, Q is the reverse complement of K, R is the third portion of the multi-branched loop, and S is the reverse complement of A. In some embodiments, the reverse complement portions are not 100% complementary, such that the stem structure may contain one or more mismatches or bulges, or non-canonical base pairing may occur. In some embodiments, a first side of the internal loop corresponding to the target binding loop comprises a nucleotide sequence (termed LTG) that is complementary to a first target site sequence of the target DNA, and a second side of the internal loop corresponding to the target binding loop comprises a second nucleotide sequence (termed RTG) that is complementary to a second target site sequence on the opposite strand of the target DNA to the first target site sequence.In some embodiments, a first side of the internal loop corresponding to the donor binding loop contains a nucleotide sequence (referred to as LDG) that is complementary to a first donor site sequence of the donor DNA, and a second side of the internal loop corresponding to the donor binding loop contains a second nucleotide sequence (referred to as RDG) that is complementary to a second donor site sequence on the opposite strand of the donor DNA relative to the first target site sequence. In some embodiments, the stem structure may contain one or more mismatches or bulges. In some embodiments, the nucleotides At least two portions of nucleotide sequence S are not complementary to the nucleotide sequence of portion A, such that the stem structure formed by base pairing between portions A and S contains two bulges. In some embodiments, at least one portion of nucleotide sequence I is not complementary to the nucleotide sequence of portion C, such that the stem structure formed by base pairing between portions I and C contains a bulge. In another embodiment, one or more nucleotides are present between the 5' end and portion A. In another embodiment, one or more nucleotides are present between the 3' end and portion S. In some embodiments, the bridge RNA binds to the IS1111_229727 transposase. In some embodiments, the donor binding loop and / or the target binding loop can form a stem or partial stem structure when not bound to the transposase, and can form a single-stranded structure when bound to the transposase.

[0187] Thus, in some embodiments, the bridge RNA comprises a nucleotide sequence comprising the following secondary structure: a first stem-multibranched loop-a second stem comprising a stem-loop and an internal loop-a third stem comprising a stem-loop and an internal loop. In some embodiments, the bridge RNA comprises additional stem-loop structures, bulges, and / or loops (see, e.g., Figure 11B). In some embodiments, the bridge RNA comprises a nucleotide sequence comprising: 5'-[A]-[B]-[C]-[D]-[E]-[F]-[G]-[H]-[I]-[J]-[K]-[L]-[M]-[N]-[O]-[P]-[Q]-[R]-[S]-3', where A is a first stem portion, B is a first portion of a multi-branched loop, C is a second stem portion, D is a first side of an internal loop corresponding to the donor binding loop, E is a third stem portion, F is a first loop portion, and G is a is the reverse complement of E, H is the second side of the internal loop corresponding to the donor binding loop, I is the reverse complement of C, J is the second portion of the multi-branched loop, K is the fourth stem portion, L is the first side of the internal loop corresponding to the target binding loop, M is the fifth stem portion, N is the second loop portion, O is the reverse complement of M, P is the second side of the internal loop corresponding to the target binding loop, Q is the reverse complement of K, R is the third portion of the multi-branched loop, and S is the reverse complement of A. In some embodiments, the reverse complement portions are not 100% complementary, such that the stem structure may contain one or more mismatches or bulges, or non-canonical base pairing may occur. In some embodiments, a first side of the internal loop corresponding to the target binding loop comprises a nucleotide sequence (termed LTG) that is complementary to a first target site sequence of the target DNA, and a second side of the internal loop corresponding to the target binding loop comprises a second nucleotide sequence (termed RTG) that is complementary to a second target site sequence on the opposite strand of the target DNA to the first target site sequence.In some embodiments, a first side of the internal loop corresponding to the donor binding loop comprises a nucleotide sequence (referred to as LDG) that is complementary to a first donor site sequence of the donor DNA, and a second side of the internal loop corresponding to the donor binding loop comprises a second nucleotide sequence (referred to as RDG) that is complementary to a second donor site sequence on the opposite strand of the donor DNA relative to the first target site sequence. In some embodiments, the stem structure may comprise one or more mismatches or bulges. In some embodiments, at least two portions of nucleotide sequence S are not complementary to the nucleotide sequence of portion A, such that the stem structure formed by base pairing between portions A and S comprises two bulges. In some embodiments, at least one portion of nucleotide sequence I is not complementary to the nucleotide sequence of portion C, such that the stem structure formed by base pairing between portions I and C comprises a bulge. In another embodiment, one or more nucleotides are present between the 5' end and portion A. In another embodiment, one or more nucleotides are present between the 3' end and portion S. In some embodiments, the bridge RNA binds to the IS1111_229727 transposase. In some embodiments, the donor binding loop and / or the target binding loop can form a stem or partial stem structure when not bound to the transposase, and form a single-stranded structure when bound to the transposase. possible.

[0188] With respect to the bridge RNA described above, in some embodiments, the LTG and RTG sequences are fully complementary to their respective target site sequences in the target DNA. In some embodiments, the LTG and / or RTG sequences are partially complementary to their respective target site sequences in the target DNA, i.e., there is tolerance for non-canonical base pairing, mismatches, and / or non-contiguous sequences. In some embodiments, there is a single mismatch in the LTG or RTG nucleotide sequence. In some embodiments, there are two, three, or four mismatches, which may be within LTG, within RTG, or may span between LTG and RTG. In some embodiments, the two, three, or four mismatches are contiguous. In some embodiments, there is a single non-canonical base pair in the LTG or RTG nucleotide sequence. In some embodiments, there are two, three, or four non-canonical base pairs, which may be within LTG, within RTG, or may span between LTG and RTG. In some embodiments, the two, three, or four non-canonical base pairs are contiguous. In some embodiments, the two, three, or four non-canonical base pairs are non-contiguous. In some embodiments, the LDG and RDG nucleotide sequences are fully complementary to the respective donor site sequences of the donor DNA. In some embodiments, the LDG and / or RDG nucleotide sequences are partially complementary to the respective donor site sequences of the donor DNA, i.e., there is tolerance for non-canonical base pairing, mismatches, and / or non-contiguous sequences. In some embodiments, there is a single mismatch in the LDG or RDG nucleotide sequence. In some embodiments, there are two, three, or four mismatches, which may be within the LDG, within the RDG, or may span between the LDG and RDG. In some embodiments, the two, three, or four mismatches are contiguous. In some embodiments, two mismatches are non-contiguous. In some embodiments, there is a single non-canonical base pair in the LDG or RDG nucleotide sequence. In some embodiments, there are two, three, or four non-canonical base pairs, which may be within the LDG, within the RDG, or may span between the LDG and RDG.In some embodiments, the 2, 3, or 4 non-canonical base pairs are contiguous. In some embodiments, the 2, 3, or 4 non-canonical base pairs are non-contiguous. In some embodiments, the non-canonical base pairing is non-Watson-Crick base pairing (i.e., not a GC base pair or an AT / U base pair). In some embodiments, the non-canonical base pairing is wobble base pairing. In some embodiments, the non-canonical base pairing is Hoogsteen base pairing. In some embodiments, the non-canonical base pairing comprises an rG-dT base pair, an rU-dG base pair, an rA-dC base pair, an rC-dA base pair, an rA-dG base pair, or an rG-dG base pair.

[0189] In certain aspects, described herein are bridge RNAs that comprise a means for directing an IS110 transposase to a polynucleotide comprising a donor or target site sequence. In some embodiments, the means for directing an IS110 transposase to a polynucleotide comprising a donor or target site sequence comprises the bridge RNA sequence of Figure 15 (SEQ ID NOS: 1-348) or SEQ ID NOS: 349-10175.

[0190] Modification of the bridge RNA sequence can be used to alter the binding specificity between the bridge RNA and the IS110 transposase, for example, to increase the affinity of the bridge RNA for the IS110 transposase or IS110 transpososome.

[0191] In some embodiments, one or more bridge RNAs are transcribed from a non-coding end of the IS110 element, for example, from the integration element (core(if present)-LE-transposase-RE-core(if present)) or from a circular form of the IS110 element (transposase-RE-core(if present)-LE). In some embodiments, one or more bridge RNAs are transcribed from an engineered construct derived from the non-coding end of the IS110 element. In some embodiments, one or more bridge RNAs are transcribed from an engineered construct derived from the non-coding end of the IS110 element and the 5' or 3' portion of the CDS. In some embodiments, one or more bridge RNAs are transcribed from a nucleotide sequence comprising an RE-core-LE. In some embodiments, one or more bridge RNAs are transcribed from a nucleotide sequence comprising an RE-core-LE, where the RE-core-LE comprises an LE, core, and RE set forth in Figure 15 (SEQ ID NOs: 1-348), Figure 17 (SEQ ID NOs: 30354-30529), or SEQ ID NOs: 349-10175 or 30530-40356. In some embodiments, one or more bridge RNAs are transcribed from a nucleotide sequence comprising a portion of the RE-core-LE sequence. In some embodiments, one or more bridge RNAs are transcribed from a nucleotide sequence comprising an RE-LE. In some embodiments, one or more bridge RNAs are transcribed from a nucleotide sequence comprising a RE-LE, where the RE-LE comprises an LE and an RE as provided in Figure 15 (SEQ ID NOs: 1-348), Figure 17 (SEQ ID NOs: 30354-30529), or SEQ ID NOs: 349-10175 or 30530-40356. In some embodiments, one or more bridge RNAs are transcribed from a nucleotide sequence comprising a portion of a RE-LE sequence. In some embodiments, one or more bridge RNAs are transcribed from a nucleotide sequence comprising an LE. In some embodiments, the LE comprises an LE as provided in Figure 15 (SEQ ID NOs: 1-348), Figure 17 (SEQ ID NOs: 30354-30529), or SEQ ID NOs: 349-10175 or 30530-40356. In some embodiments, one or more bridge RNAs are transcribed from a nucleotide sequence comprising a portion of an LE sequence. In some embodiments, one or more bridge RNAs are transcribed from a nucleotide sequence comprising an RE.In some embodiments, the RE comprises an RE provided in Figure 15 (SEQ ID NOs: 1-348), Figure 17 (SEQ ID NOs: 30354-30529), or SEQ ID NOs: 349-10175 or 30530-40356. In some embodiments, one or more bridge RNAs are transcribed from a nucleotide sequence that includes a portion of the RE sequence.

[0192] Any bridge RNA described herein may include additional nucleotide moieties, nucleotide linkers, loops, and modifications (e.g., structured RNA pseudoknots or other RNA structures), as will be understood by one of skill in the art.

[0193] Any bridge RNA described herein can be expressed from linear or circular dsDNA or ssDNA, or can be supplied as synthetic RNA. When the bridge RNA is encoded by an expression construct, the bridge RNA coding sequence can be expressed from a suitable promoter, such as a promoter derived from the non-coding end of an IS110 member, or from an ectopic promoter, such as a type III pol III promoter.

[0194] In one embodiment, the present invention contemplates an approach for engineering IS110 systems. Given the abundance and diversity of IS110 orthologs, extracting different domains or sequences from closely or distantly related orthologs and creating new combinations can lead to new IS110 systems with advantageous enzymatic properties. In one embodiment, one IS110 ortholog may have a highly efficient, specific RuvC-like DEDD catalytic domain, while another IS110 may have a bridge RNA with a highly specific target and / or donor loop. These IS110 systems can be combined to exploit their respective favorable properties. In another embodiment, the bridge RNA-binding subdomain and transposase domain of the RuvC-like DEDD catalytic domain may have affinity for a specific bridge RNA, which can be swapped and combined with related systems to exploit their respective favorable properties. In another embodiment, two or more different bridge RNA structures can be compared, and their subsequences and domains can be swapped to exploit favorable properties. This may result in new bridge RNA species with different properties.

[0195] For example, we can model the bridge RNA structure using the Infernal (Nawrocki and Models can be constructed using standard bioinformatics software packages such as (Eddy 2013). These models can be searched against genome sequence databases to identify bridge RNA sequences, and the lengths and sequences of the predicted guide subsequences of the target and donor site binding loops for each candidate can be calculated. Candidate IS110 systems with particularly long guides encoded within the target and / or donor binding loops can be prioritized for experimental characterization. In another embodiment, the target loop sequence can be searched against nearby flanking sequences to identify complementary target sequences, and candidates with particularly long matches between the target site and the target binding loop can be prioritized for experimental characterization. In another embodiment, the bridge RNA donor binding loop sequence can be searched against IS110 end sequences to identify complementary donor sequences, and candidates with particularly long matches between the target site and the bridge RNA target binding loop can be prioritized for experimental characterization.

[0196] E. Donor and Target Sites The donor and target site sequences can be any sequences that can be bound by a transposase in conjunction with a bridge RNA. For example, the bridge RNA target binding loop encodes LTG and RTG, which have base-pairing specificity for the target site sequences within the target molecule, i.e., the left target (LT) and right target (RT) sequences. In some embodiments, the bridge RNA donor binding loop encodes LDG and RDG, which have base-pairing specificity for the donor site sequences within the donor molecule, i.e., the left donor (LD) and right donor (RD) sequences, thereby positioning the transposase to mediate site-specific transposition between DNA molecules containing the target site and the donor site. The target site sequence and the donor site sequence can be on the same or different molecules. The bridge RNA binding loop can also bind to the core sequence of the target site or donor site sequence. As noted above, certain mismatches and / or discontinuities can be tolerated between the LTG, RTG, and target site sequences, and / or between the LDG, RDG, and donor site sequences.

[0197] Described herein is the use of IS110 family transposases to mediate insertion, excision recombination, and inversion reactions. The type of reaction mediated by an IS110 family transposase is a function of the orientation and DNA strand location of the LT and RT target site sequences and the LD and RD donor site sequences, as described below.

[0198] Described herein is the use of a wild-type bridge RNA, or a bridge RNA comprising a wild-type donor-binding loop and / or a target-binding loop, to mediate recombination between DNA sequences comprising wild-type target and donor site sequences. For example, when a donor site sequence of any of the IS110 family transposases described herein (see, e.g., Figure 17 (SEQ ID NOS: 30354-30529) or SEQ ID NOS: 30530-40356) is present on a DNA molecule, and the corresponding target site sequence (see, e.g., Figure 18 (SEQ ID NOS: 20351-20526) or SEQ ID NOS: 20527-30353) is present on a DNA molecule, expression of the IS110 family transposase and the corresponding bridge RNA can result in insertion, excision recombination, or inversion.

[0199] In some embodiments, a wild-type bridge RNA, or a bridge RNA comprising a wild-type donor binding loop and a target binding loop, may be used to mediate an insertion, excision recombination, or inversion reaction between a DNA sequence comprising a wild-type target site sequence and a DNA sequence comprising a wild-type donor site sequence.

[0200] Also described herein are methods for mediating transposition between DNA sequences of interest. Reprogramming of the donor RNA binding loop and / or target binding loop sequences.

[0201] In some embodiments, a bridge RNA comprising a wild-type donor binding loop may be used in combination with a reprogrammed target binding loop to mediate an insertion, excision recombination, or inversion reaction between a DNA sequence comprising a wild-type donor site sequence and a DNA sequence comprising a target site sequence of interest. In some embodiments, a bridge RNA comprising a reprogrammed donor binding loop may be used in combination with a wild-type target binding loop to mediate an insertion reaction between a DNA sequence comprising a donor site sequence of interest and a wild-type target site sequence.

[0202] In some embodiments, a bridge RNA comprising a reprogrammed donor binding loop may be used in combination with a reprogrammed target binding loop to mediate an insertion reaction between a DNA sequence comprising a donor site sequence of interest and a DNA sequence comprising a target site sequence of interest.

[0203] In some embodiments, different bridge RNAs (in combination with the same or different transposases) may be used to simultaneously target multiple target site sequences or donor site sequences.

[0204] E.1. Insertion reaction To mediate the insertion reaction, the target site and donor site sequences are present on separate DNA molecules. When the donor site sequence is located on a circular DNA molecule, the transposition reaction functionally results in the insertion of a circular DNA molecule containing the donor site sequence into the target site sequence of a second DNA molecule. When the target site sequence is located on a circular DNA molecule, the transposition reaction functionally results in the insertion of a circular DNA molecule containing the target site sequence into the donor site sequence of a second DNA molecule. When both the donor site sequence and the target site sequence are located on a linear DNA molecule, the transposition reaction functionally results in the insertion of a linear DNA molecule containing the donor site sequence into the target site sequence of a second DNA molecule, and the double-strand break that occurs after recombination with the linear DNA molecule is repaired by endogenous DNA repair pathways such as non-homologous end joining. When both the donor site sequence and the target site sequence are located on different linear chromosomes, the transposition reaction functionally results in a chromosomal translocation without a double-strand break.

[0205] In embodiments where a core sequence is absent, the sequence of LTG in the target binding loop in the 5' to 3' direction is complementary to the first strand of the target site sequence to mediate insertion. The sequence of RTG in the target binding loop in the 5' to 3' direction is reverse complementary to the opposite strand of the first strand of the target site sequence, with the 3' end of RTG being reverse complementary to the nucleotide immediately following the target site sequence complementary to LTG. In some embodiments, the sequence of RTG in the target binding loop in the 5' to 3' direction is reverse complementary to the opposite strand of the first strand of the target site sequence, with the 3' end of RTG being reverse complementary to the second or third nucleotide immediately following the target site sequence complementary to LTG (i.e., there is a 1-2 nucleotide gap between the LTG and RTG binding sites). The sequence of LDG in the donor binding loop in the 5' to 3' direction is complementary to the first strand of the donor site sequence. The 5' to 3' sequence of RDG is the reverse complement of the first strand of the donor site sequence, and the 3' end of RDG is the reverse complement of the nucleotide immediately following the donor site sequence complementary to LDG. In some embodiments, the 5' to 3' sequence of RDG is the reverse complement of the first strand of the donor site sequence, and the 3' end of RDG is the reverse complement of the second or third nucleotide immediately following the donor site sequence complementary to LDG (i.e., there is a 1-2 nucleotide gap between the LDG and RDG bond).

[0206] In some embodiments where a core sequence is present, a 5' to 3' sequence is required to mediate insertion. The LTG sequence of the target binding loop in the forward direction is complementary to the first strand of the target site sequence, and the 3' end of the LTG is complementary to at least one nucleotide of the core sequence on the first strand of the target site sequence. The RTG sequence in the 5' to 3' direction is reverse complementary to the opposite strand of the target site sequence, and the 3' end of the RTG is reverse complementary to at least one nucleotide of the core sequence on the opposite strand of the target site sequence. In some embodiments, there is a 1-2 nucleotide gap between the LTG and RTG binding sites. The LDG sequence in the 5' to 3' direction is complementary to the first strand of the donor site sequence, and the 3' end of the LDG is complementary to at least one nucleotide of the core sequence on the first strand of the donor site sequence. The 5' to 3' sequence of the RDG is the reverse complement of the strand opposite the first strand of the donor site sequence, and the 3' end of the RDG is reverse complementary to at least one nucleotide of the core sequence on the strand opposite the first strand of the donor site sequence. In some embodiments, there is a 1-2 nucleotide gap between the binding sites of the LDG and RDG. See Figure 4, Figure 35B.

[0207] In some embodiments where a core sequence is present, to mediate insertion, the sequence of LTG in the target binding loop in the 5' to 3' direction is complementary to the first strand of the target site sequence. The sequence of RTG in the 5' to 3' direction is the reverse complementary sequence to the opposite strand of the first strand of the target site sequence. The sequence of LDG in the 5' to 3' direction is complementary to the first strand of the donor site sequence. The sequence of RDG in the 5' to 3' direction is the reverse complementary sequence to the opposite strand of the first strand of the donor site sequence. Thus, in this embodiment, RTG, LTG, RDG, and LDG do not bind to the core sequence, even if present. In some embodiments, one or more of RTG, LTG, RDG, and LDG can be complementary to at least one nucleotide of the core sequence. See, e.g., Figure 35B. Furthermore, in some embodiments, there is a 1-2 nucleotide gap between where LDG and RDG bind and / or where LTG and RTG bind. In some embodiments, the bridge RNA is engineered or modified so that one or more of RTG, LTG, RDG, and LDG no longer bind to the core sequence, even though they are present in the donor and target site sequences. For example, a naturally occurring bridge RNA in which one or more of RTG, LTG, RDG, and LDG bind to the core sequence can be modified so that one or more of RTG, LTG, RDG, and LDG no longer bind to the core sequence, even though they remain present in the donor and target site sequences. Such an approach can increase the binding specificity of the bridge RNA for the donor site sequence and / or the target site sequence. See, e.g., Figures 40A-D.

[0208] In any of the above embodiments, the LTG and RTG sequences are fully complementary to their respective target site sequences in the target DNA. In some embodiments, the LTG and / or RTG are partially complementary to their respective target site sequences in the target DNA, i.e., there is tolerance for non-canonical base pairing, mismatches, and / or non-contiguous sequences. In some embodiments, there is a single mismatch in the LTG or RTG nucleotide sequence. In some embodiments, there are two, three, or four mismatches, which may be within LTG, within RTG, or may span LTG and RTG. In some embodiments, the two, three, or four mismatches are contiguous. In some embodiments, there is a single non-canonical base pair in the LTG or RTG nucleotide sequence. In some embodiments, there are two, three, or four non-canonical base pairs, which may be within LTG, within RTG, or may span LTG and RTG. In some embodiments, the two, three, or four non-canonical base pairs are contiguous. In some embodiments, the 2, 3, or 4 non-canonical base pairs are non-contiguous. In some embodiments, the LDG and RDG nucleotide sequences are fully complementary to the respective donor site sequences of the donor DNA. In some embodiments, the LDG and / or RDG nucleotide sequences are partially complementary to the respective donor site sequences of the donor DNA, i.e., That is, there is tolerance for non-canonical base pairing, mismatches, and / or non-contiguous base pairing. In some embodiments, there is a single mismatch in the LDG or RDG nucleotide sequence. In some embodiments, there are 2, 3, or 4 mismatches, which may be within the LDG, within the RDG, or may span across the LDG and RDG. In some embodiments, the 2, 3, or 4 mismatches are contiguous. In some embodiments, the 2, 3, or 4 mismatches are non-contiguous. In some embodiments, there is a single non-canonical base pair in the LDG or RDG nucleotide sequence. In some embodiments, there are 2, 3, or 4 non-canonical base pairs, which may be within the LDG, within the RDG, or may span across the LDG and RDG. In some embodiments, the 2, 3, or 4 non-canonical base pairs are contiguous. In some embodiments, the 2, 3, or 4 non-canonical base pairs are non-contiguous. In some embodiments, the non-canonical base pairing is non-Watson-Crick base pairing (i.e., not a GC base pair or an AT / U base pair). In some embodiments, the non-canonical base pairing is a wobble base pairing. In some embodiments, the non-canonical base pairing is a Hoogsteen base pairing. In some embodiments, the non-canonical base pairing comprises an rG-dT base pair, an rU-dG base pair, an rA-dC base pair, an rC-dA base pair, an rA-dG base pair, or an rG-dG base pair.

[0209] In some embodiments, it is not necessary to define the sequences of the donor site and / or target site, or LT, RT, LD, and RD (and thus LTG, RTG, LDG, and RDG). Alternatively, an insertion reaction can be mediated between a DNA molecule containing a donor site that includes a wild-type RE-LE or a subsequence thereof, or, if a core is used, an RE-core-LE or a subsequence thereof of any of the IS110 family transposases described herein (see Figure 15 (SEQ ID NOS: 1-348), Figure 17 (SEQ ID NOS: 30354-30529), or SEQ ID NOS: 349-10175 or 30530-40356), and a DNA molecule containing a target sequence that includes an LF-RF or a subsequence thereof, or, if a core is used, an LF-core-RF of any of the IS110 family transposases described herein (see Figure 18 (SEQ ID NOS: 20351-20526), ​​or SEQ ID NOS: 20527-30353), by providing an IS110 family transposase and its corresponding bridge RNA. Because the bridge RNA is encoded by the LE or RE, the sequence of the bridge RNA need not be defined. Thus, in some embodiments, a bridge RNA sequence may be provided by providing the nucleotide sequence of an LE or RE.

[0210] For the insertion reactions described herein, the system may further include one or more polynucleotides for insertion (e.g., for insertion into a target site) comprising a cargo and a donor site sequence. In another embodiment, the system further includes one or more polynucleotides for insertion (e.g., for insertion into a donor site) comprising a cargo sequence and a target site sequence. In some embodiments, the polynucleotide for insertion is circular. In some embodiments, the polynucleotide for insertion is linear. When the polynucleotide for insertion is linear, the double-stranded break that occurs after recombination with the linear DNA molecule is repaired by endogenous DNA repair pathways, such as non-homologous end joining.

[0211] For insertion reactions resulting in cassette exchange (e.g., recombinase-mediated cassette exchange (RMCE)), in some embodiments, the one or more polynucleotides for insertion further comprise a second donor site sequence that is different from the first donor site sequence (i.e., the one or more polynucleotides for insertion comprise two donor site sequences). In some embodiments, the one or more polynucleotides for insertion further comprise a target site sequence that corresponds to a donor site that is different from the first donor site sequence (i.e., the one or more polynucleotides for insertion comprise a donor site sequence and a target site sequence). In some embodiments, the one or more polynucleotides for insertion further comprise a second target site sequence that is different from the first target site sequence. In some embodiments, the one or more polynucleotides for insertion further comprise a donor site sequence corresponding to a target site that is different from the first target site sequence (i.e., the one or more polynucleotides for insertion comprise a donor site sequence and a target site sequence). In such embodiments, the system includes a transposase having a first bridge RNA targeted to a first donor or target site sequence and a transposase having a second bridge RNA targeted to a second donor or target site sequence. In some embodiments, the transposases bound to the first bridge RNA and the second bridge RNA are the same type, e.g., IS621. In some embodiments, the transposases bound to the first bridge RNA and the second bridge RNA are different transposases.

[0212] In some embodiments, the system includes components for carrying out at least two (and in some embodiments, two or more) desired reactions. An insertion polynucleotide, e.g., on a plasmid, including a cargo, a donor site sequence, and a target site sequence, is recombined with a bridge RNA targeting the donor and target site sequences, creating two minicircles: one with an LT-RD junction and the other with an LD-RT junction. A second bridge RNA, including a binding loop with specificity for the newly formed LT-RD or LD-RT and a binding loop targeting a site in the genome of interest, binds to the minicircle and integrates it into the target site sequence, e.g., the genome of interest. Thus, for such an insertion reaction, in some embodiments, the system can include one or more first circular polynucleotides (e.g., plasmids) including a cargo, a first donor site sequence including LD and RD sequences, and a second target site sequence including LT and RT sequences. In some embodiments, the system further includes a transposase having a first bridge RNA targeting a first donor site sequence and a first target site sequence on the first polynucleotide, and a second bridge RNA targeting a second donor site sequence and a second target site sequence, where the second donor site sequence includes the LT of the first target site sequence and the RD of the first donor site, or the LD of the first donor site and the RT of the first target site, and the second target site sequence is a site within a target (e.g., a genome) of interest. The second bridge RNA can target the second donor site sequence using the donor binding loop (i.e., the donor binding loop of the second bridge RNA targets the second donor site sequence) or the target binding loop (i.e., the target binding loop of the second bridge RNA targets the second donor site sequence), as long as the other loop targets the site (e.g., a genome) of interest. In some embodiments, the system causes recombination of the donor site sequence and the target site sequence of the first polynucleotide to produce a first minicircle comprising an LT-RD junction and a second minicircle having an LD-RT junction.Depending on the orientation of the cargo, first donor site sequence, and first target site sequence, either the first minicircle contains the cargo, or the second minicircle contains the cargo and the second bridge RNA targets either the cargo-containing minicircle. In some embodiments, the system causes recombination between either the first or second minicircle, whichever contains the cargo, and the site of interest, resulting in insertion of the cargo at the site of interest.

[0213] In some embodiments, the first circular polynucleotide further comprises a second cargo, and the system further comprises a third bridge RNA targeting a third donor site sequence and a third target site sequence, where the third donor site sequence comprises the LT of the first target site sequence and the RD of the first donor site, or the LD of the first donor site and the RT of the first target site (e.g., either the LT-RD or LD-RT that is not targeted by the second bridge RNA), and the second target site sequence is a second site within the target of interest (e.g., a genome). The third bridge RNA uses the donor binding loop (i.e., the donor binding of the third bridge RNA) to target the second site of interest (e.g., a genome) as long as the other loop targets the second site of interest (e.g., a genome). The third bridge RNA may target a third donor site sequence using a target binding loop (i.e., the loop targets the third donor site sequence) or a target binding loop (i.e., the target binding loop of the third bridge RNA targets the third donor site sequence). In some embodiments, the system causes the donor site sequence and target site sequence of the first polynucleotide to recombine, resulting in the creation of a first minicircle comprising an LT-RD junction and a first cargo and a second minicircle comprising an LD-RT junction and a second cargo, or vice versa (e.g., the second cargo is comprised on the minicircle with the LT-RD junction and the first cargo is comprised on the minicircle with the LD-RT junction). In some embodiments, the system causes the first minicircle to recombine with a first site within the target of interest and the second minicircle to recombine with a second site within the target of interest, or vice versa, resulting in the insertion of the first cargo and the second cargo into their respective sites of interest.

[0214] In some embodiments, in systems comprising one or more polynucleotides for insertion comprising a cargo and a donor site sequence, the cargo sequence is oriented in the 5' to 3' direction relative to the donor site sequence. In some embodiments, in systems comprising one or more polynucleotides for insertion comprising a cargo and a donor site sequence, the cargo sequence is oriented in the 3' to 5' direction relative to the donor site sequence. In some embodiments, in systems comprising one or more polynucleotides for insertion comprising a cargo and two donor site sequences, the cargo sequence is oriented in the 5' to 3' direction between the two donor site sequences. In some embodiments, in systems comprising one or more polynucleotides for insertion comprising a cargo and two donor site sequences, the cargo sequence is oriented in the 3' to 5' direction between the two donor site sequences. In some embodiments, in systems comprising one or more polynucleotides for insertion comprising a cargo, a donor site sequence, and target site sequences corresponding to different donor sites, the cargo sequence is oriented in the 5' to 3' direction between the donor site sequence and the target site sequence. In some embodiments, in systems comprising one or more polynucleotides for insertion comprising a cargo, a donor site sequence, and a target site sequence corresponding to a different donor site, the cargo sequence is oriented in the 3' to 5' direction between the donor site sequence and the target site sequence.

[0215] In some embodiments, in systems comprising one or more polynucleotides for insertion comprising a cargo and a target site sequence, the cargo sequence is oriented in the 5' to 3' direction relative to the target site sequence. In some embodiments, in systems comprising one or more polynucleotides for insertion comprising a cargo and a target site sequence, the cargo sequence is oriented in the 3' to 5' direction relative to the target site sequence. In some embodiments, in systems comprising one or more polynucleotides for insertion comprising a cargo and two target site sequences, the cargo sequence is oriented in the 5' to 3' direction between the two target site sequences. In some embodiments, in systems comprising one or more polynucleotides for insertion comprising a cargo and two target site sequences, the cargo sequence is oriented in the 3' to 5' direction between the two target site sequences. In some embodiments, in systems comprising one or more polynucleotides for insertion comprising a cargo, a target site sequence, and a donor site sequence corresponding to a different target site, the cargo sequence is oriented in the 5' to 3' direction between the target site sequence and the donor site sequence. In some embodiments, in systems comprising one or more polynucleotides for insertion comprising a cargo, a target site sequence, and a donor site sequence, the cargo sequence is oriented in the 3' to 5' direction between the target site sequence and the donor site sequence.

[0216] The insertion polynucleotide may be equivalent to a transposable element that can be inserted or integrated into a target site sequence or a donor site sequence. The insertion polynucleotide may be or include one or more components of a transposon.

[0217] The cargo of the insert polynucleotide may be a gene, a gene fragment, or a non-coding polynucleotide. The polynucleotide may comprise any type of polynucleotide, including but not limited to a nucleic acid sequence, a regulatory polynucleotide, or a synthetic polynucleotide.

[0218] The insertion polynucleotide may comprise a transposon left end (LE) and a transposon right end (RE). The LE or RE sequence may be an endogenous sequence of the IS110 used, or a heterologous sequence recognizable by the IS110 used, or a synthetic sequence containing sequence or structural features recognized by IS110 and sufficient to allow insertion of the polynucleotide into the target site sequence. In certain exemplary embodiments, the LE or RE sequence is truncated. In certain exemplary embodiments, the LE or RE sequence is 20 to 500 base pairs, 500 to 490 base pairs, 500 to 480 base pairs, 500 to 470 base pairs, 500 to 460 base pairs, 500 to 450 base pairs, 500 to 440 base pairs, 500 to 430 base pairs, 500 to 420 base pairs, 500 to 410 base pairs, 500 to 400 base pairs, 400 to 500 base pairs, or a combination thereof. 390 base pairs, 400-380 base pairs, 400-370 base pairs, 400-360 base pairs, 400-350 base pairs, 400-340 base pairs, 400-330 base pairs, 400-320 base pairs, 400-310 base pairs, 400-300 base pairs, 300-290 base pairs, 300-280 base pairs, 300-270 base pairs, 300-260 base pairs, 30 0-250 base pairs, 300-240 base pairs, 300-230 base pairs, 300-220 base pairs, 300-210 base pairs, 300-200 base pairs, 200-100 base pairs, 100-190 base pairs, 100-180 base pairs, 100-170 base pairs, 100-160 base pairs, 100-150 base pairs, 100-140 base pairs, 100-130 base pairs, The length of the polynucleotide may be 100 to 120 base pairs, 100 to 110 base pairs, 20 to 100 base pairs, 20 to 90 base pairs, 20 to 80 base pairs, 20 to 70 base pairs, 20 to 60 base pairs, 20 to 50 base pairs, 20 to 40 base pairs, 20 to 30 base pairs, 50 to 100 base pairs, 60 to 100 base pairs, 70 to 100 base pairs, 80 to 100 base pairs, or 90 to 100 base pairs. In some embodiments, the polynucleotide for insertion may comprise transposon LD and transposon RD.

[0219] The polynucleotide for insertion can comprise a transposon left flank (LF) and a transposon right flank (RF). The LF or RF sequence can be an endogenous sequence of the IS110 used, or a heterologous sequence recognizable by the IS110 used, or a synthetic sequence containing sequence or structural features recognized by IS110 and sufficient to allow insertion of the polynucleotide into the donor site. In certain exemplary embodiments, the LF or RF sequence is truncated. In certain exemplary embodiments, the LF or RF sequence is 20 to 500 base pairs, 500 to 490 base pairs, 500 to 480 base pairs, 500 to 470 base pairs, 500 to 460 base pairs, 500 to 450 base pairs, 500 to 440 base pairs, 500 to 430 base pairs, 500 to 420 base pairs, 500 to 410 base pairs, 500 to 400 base pairs, 400 to 500 base pairs, or a similar sequence. 390 base pairs, 400-380 base pairs, 400-370 base pairs, 400-360 base pairs, 400-350 base pairs, 400-340 base pairs, 400-330 base pairs, 400-320 base pairs, 400-310 base pairs, 400-300 base pairs, 300-290 base pairs, 300-280 base pairs, 300-270 base pairs, 300-260 base pairs, 30 0-250 base pairs, 300-240 base pairs, 300-230 base pairs, 300-220 base pairs, 300-210 base pairs, 300-200 base pairs, 200-100 base pairs, 100-190 base pairs, 100-180 base pairs, 100-170 base pairs, 100-160 base pairs, 100-150 base pairs, 100-140 base pairs, 100-130 base pairs, The length may be 100 to 120 base pairs, 100 to 110 base pairs, 20 to 100 base pairs, 20 to 90 base pairs, 20 to 80 base pairs, 20 to 70 base pairs, 20 to 60 base pairs, 20 to 50 base pairs, 20 to 40 base pairs, 20 to 30 base pairs, 50 to 100 base pairs, 60 to 100 base pairs, 70 to 100 base pairs, 80 to 100 base pairs, or 90 to 100 base pairs. In some embodiments, the polynucleotide for insertion may comprise a transposon LT and a transposon RT.

[0220] As used herein, the terms "cargo," "cargo gene," or "cargo sequence" refer to a gene or nucleic acid sequence that can be integrated into a target site sequence or a donor site sequence via transposition. The terms "delivered cargo," "delivered cargo gene," or "delivered cargo sequence" refer to any gene, genetic system, regulatory sequence, or sequence that can be delivered and integrated into a target site sequence or a donor site sequence via a transposition event. In some embodiments, the cargo gene or sequence is delivered to a target cell in vitro, in vivo, or ex vivo.

[0221] In some embodiments, the cargo gene or sequence delivered is a biologically active agent, i.e., one that has activity in a cell, organ, tissue, and / or subject. For example, a gene or sequence that, when administered to a subject, has a biological effect on the subject, is considered to be biologically active. In some embodiments, the cargo gene or sequence delivered is a therapeutic agent. As used herein, the term "therapeutic agent" refers to any agent that has a beneficial effect when administered to a subject. In some embodiments, the cargo gene or sequence delivered to a cell is a transcription factor, tumor suppressor, developmental regulator, growth factor, metastasis suppressor, pro-apoptotic protein, nuclease, or recombinase. In some embodiments, the cargo gene or sequence encodes a protein, some non-limiting examples of which include p53, Rb (retinoblastoma protein), BRCA1, BRCA2, PTEN, APC, CD95, ST7, ST14, BCL-2 family proteins, caspases; BRMS1, CRSP3, DRG1, KAI1, KISS1, NM23, TIMP family proteins, BMP family growth factors, EGF, EPO, FGF, G-CSF, GM-CSF, GDF family growth factors, HGF, HDGF, IGF, PDGF, TPO, TGF-α, TGF-β, VEGF; zinc finger nucleases, Cre, Dre, or FLP recombinase. In some embodiments, the cargo gene or sequence is associated with a small molecule. In some embodiments, the cargo gene or sequence to be delivered is a diagnostic agent. In some embodiments, the cargo gene or sequence to be delivered is a prophylactic agent. In some embodiments, the cargo gene or sequence to be delivered is useful as an imaging agent. In some embodiments, the diagnostic or imaging agent is biologically active, and in other embodiments, it is not.

[0222] In some embodiments, the insert polynucleotide is from 11 bases (b) or base pairs (bp) to about 100 kilobases (kb) or kilobase pairs (kbp) or more in length (e.g., about 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 20, 21, 22, 23, 2 From 5 or 100 b or bp to about 110, 120, 125, 150, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 525, 550, 575, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1250, 1500, 1750, 2000, 2250, 2500, 2750, 3000, 3250, 350 0, 3750, 4000, 4250, 4500, 4750, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10,000, 10,500, 11,000, 11,500, 12,000, 12,500, 13,000, 13,500, 14,000, 14,500, 15,000, 16,000, 17,000, 18,000, 19,000 , 20,000, 21,000, 22,000, 23,000, 24,000, 25,000, 26,000, 27,000, 28,000, 29,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, or 100,000 b or bp in length), with the upper limit corresponding to the delivery limit for delivering the polynucleotide for insertion into the cells of interest.

[0223] Inserting polynucleotides can be delivered as dsDNA, or, if the cellular machinery or additional components are delivered to make these molecules dsDNA, as ssDNA or RNA. Polynucleotides can be provided in the form of circular or linear plasmids, or as components of vectors (e.g., as components of viral vectors), or as amplification or polymerization products thereof. Shorter DNA molecules can be provided as double-stranded oligonucleotides. Exemplary double-stranded template oligonucleotides are 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 95, 96, 97, 98, 99, 100, 110, 115, 120, 125, 150, 175, 200, 225, or 250 b or bp, or at least more. The inserting polynucleotide may be provided in the reaction mixture for introduction into the cell at a concentration of about 1 μM to about 200 μM, about 2 μM to about 190 μM, about 2 μM to about 180 μM, about 5 μM to about 180 μM, about 9 μM to about 180 μM, about 10 μM to about 150 μM, about 20 μM to about 140 μM, about 30 μM to about 130 μM, about 40 μM to about 120 μM, or about 45 or 50 μM to about 90 or 100 μM.In some cases, the donor DNA is at 1 μM, 2 μM, 3 μM, 4 μM, 5 μM, 6 μM, 7 μM, 8 μM, 9 μM, 10 μM, 11 μM, 12 μM, 13 μM, 14 μM, 15 μM, 16 μM, 17 μM, 18 μM, 19 μM, 20 μM, 25 μM, 30 μM, 35 μM, 40 μM, 45 ...5 μM, 6 μM, 7 μM, 8 μM, 9 μM, 10 μM, 11 μM, 12 μM, 13 μM, 14 μM, 15 μM, 16 μM, 17 μM, 18 μM, 19 μM, 20 μM, 25 μM, 30 μM, 35 μM, 40 μ The antibody may be provided in the reaction mixture at a concentration of 0 μM, 55 μM, 60 μM, 70 μM, 80 μM, 90 μM, 100 μM, 110 μM, 115 μM, 120 μM, 130 μM, 140 μM, 150 μM, 160 μM, 170 μM, 180 μM, 190 μM, 200 μM, or more, or at about or greater than that concentration.

[0224] The polynucleotide for insertion can comprise a wide variety of different sequences. In some cases, the polynucleotide encodes a stop codon or a frameshift compared to the target genomic region before insertion. Such polynucleotides can be useful for knocking out or inactivating a gene or a portion thereof. In some cases, the polynucleotide encodes one or more missense mutations or in-frame insertions or deletions compared to the target genomic region. Such polynucleotides can be useful for altering the expression level or activity (e.g., ligand specificity) of a target gene or a portion thereof.

[0225] As described above, the polynucleotide for insertion comprises a donor site sequence and / or a target site sequence for insertion into a DNA sequence comprising the target site sequence and / or the donor site sequence, respectively. The target site sequence and / or the donor site sequence to be inserted can be located on any polynucleotide sequence of interest, including, but not limited to, genomic DNA and plasmids. In some embodiments, the target site sequence and / or the donor site sequence to be inserted is a polynucleotide sequence present in the genome or DNA of interest. In some embodiments, the target site sequence and / or the donor site sequence is naturally present in the genome or DNA of interest. In some embodiments, the target site sequence and / or the donor site sequence to be inserted is introduced into the genome or DNA of interest. Methods for introducing DNA sequences, such as target site sequences or donor site sequences, into a genome or DNA of interest are known in the art and include, but are not limited to, CRISPR-Cas9, homology-directed repair (HDR), transposase, integrase, etc. In some embodiments, the genomic DNA is located within a cell. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a prokaryotic organism. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a mouse cell. In some embodiments, the cell is a stem cell.

[0226] The target site sequence for insertion can include a transposon left flank (LF) and a transposon right flank (RF). The LF or RF sequence can be an endogenous sequence of the IS110 used, or a heterologous sequence recognizable by the IS110 used, or a synthetic sequence containing sequence or structural features recognized by IS110 and sufficient to allow insertion of the polynucleotide into the donor site sequence. In certain exemplary embodiments, the LF or RF sequence is truncated. In certain exemplary embodiments, the LF or RF sequence is 20 to 500 base pairs, 500 to 490 base pairs, 500 to 480 base pairs, 500 to 470 base pairs, 500 to 460 base pairs, 500 to 450 base pairs, 500 to 440 base pairs, 500 to 430 base pairs, 500 to 420 base pairs, 500 to 410 base pairs, 500 to 400 base pairs, 400 to 500 base pairs, or a combination thereof. 390 base pairs, 400-380 base pairs, 400-370 base pairs, 400-360 base pairs, 400-350 base pairs, 400-340 base pairs, 400-330 base pairs, 400-320 base pairs, 400-310 base pairs, 400-300 base pairs, 300-290 base pairs, 300-280 base pairs, 300-270 base pairs, 300-260 base pairs, 30 0-250 base pairs, 300-240 base pairs, 300-230 base pairs, 300-220 base pairs, 300-210 base pairs, 300-200 base pairs, 200-100 base pairs, 100-190 base pairs, 100-180 base pairs, 100-170 base pairs, 100-160 base pairs, 100-150 base pairs, 100-140 base pairs, 100-130 base pairs, The length may be 100 to 120 base pairs, 100 to 110 base pairs, 20 to 100 base pairs, 20 to 90 base pairs, 20 to 80 base pairs, 20 to 70 base pairs, 20 to 60 base pairs, 20 to 50 base pairs, 20 to 40 base pairs, 20 to 30 base pairs, 50 to 100 base pairs, 60 to 100 base pairs, 70 to 100 base pairs, 80 to 100 base pairs, or 90 to 100 base pairs. In some embodiments, the target site sequence to be inserted may include a transposon left target (LT) and a transposon right target (RT).

[0227] The donor site sequence to be inserted can include a transposon left end (LE) and a transposon right end (RE). The LE or RE sequence can be an endogenous sequence of the IS110 used, or a heterologous sequence recognizable by the IS110 used, or a synthetic sequence containing sequence or structural features recognized by IS110 and sufficient to allow insertion of a polynucleotide into the donor site sequence. In certain exemplary embodiments, the LE or RE sequence is truncated. In certain exemplary embodiments, the LE or RE sequence is 20 to 500 base pairs, 500 to 490 base pairs, 500 to 480 base pairs, 500 to 470 base pairs, 500 to 460 base pairs, 500 to 450 base pairs, 500 to 440 base pairs, 500 to 430 base pairs, 500 to 420 base pairs, 500 to 410 base pairs, 500 to 400 base pairs, 400 to 500 base pairs, or a combination thereof. 390 base pairs, 400-380 base pairs, 400-370 base pairs, 400-360 base pairs, 400-350 base pairs, 400-340 base pairs, 400-330 base pairs, 400-320 base pairs, 400-310 base pairs, 400-300 base pairs, 300-290 base pairs, 300-280 base pairs, 300-270 base pairs, 300-260 base pairs, 30 0-250 base pairs, 300-240 base pairs, 300-230 base pairs, 300-220 base pairs, 300-210 base pairs, 300-200 base pairs, 200-100 base pairs, 100-190 base pairs, 100-180 base pairs, 100-170 base pairs, 100-160 base pairs, 100-150 base pairs, 100-140 base pairs, 100-130 base pairs, The length may be 100 to 120 base pairs, 100 to 110 base pairs, 20 to 100 base pairs, 20 to 90 base pairs, 20 to 80 base pairs, 20 to 70 base pairs, 20 to 60 base pairs, 20 to 50 base pairs, 20 to 40 base pairs, 20 to 30 base pairs, 50 to 100 base pairs, 60 to 100 base pairs, 70 to 100 base pairs, 80 to 100 base pairs, or 90 to 100 base pairs. The donor site sequence to be inserted may include transposon LD and transposon RD.

[0228] E.2. Excision recombination and inversion To mediate excision recombination or inversion reactions, the target site sequence and donor site sequence are present on a single DNA molecule. When the target site sequence and donor site sequence are present on a single DNA molecule and are arranged such that the LT and LD are on the same DNA strand and the RT and RD are on the same DNA strand, the insertion reaction functionally results in excision of the DNA sequence intervening between the target site and donor site (excision recombination). When the target site sequence and donor site sequence are present on a single DNA molecule and are arranged such that the LT and RD are on the same DNA strand and the RT and LD are on the same DNA strand, the insertion reaction functionally results in inversion of the DNA sequence intervening between the target site and donor site. In some embodiments where a core sequence is present, the core sequence of the donor site sequence and the core sequence of the target site sequence are on opposite strands. An inversion reaction on a chromosomal scale can be referred to as an intrachromosomal translocation.

[0229] In some embodiments, to mediate excision recombination in the absence of a core sequence, the 5' to 3' sequence of the LTG in the target binding loop is complementary to the first strand of the target site sequence. The 5' to 3' sequence of the RTG in the target binding loop is reverse complementary to the opposite strand of the target site sequence from the first strand, with the 3' end of the RTG being reverse complementary to the nucleotide immediately following the target site sequence complementary to LTG. In some embodiments, the 5' to 3' sequence of the RTG in the target binding loop is reverse complementary to the opposite strand of the target site sequence from the first strand, with the 3' end of the RTG being reverse complementary to the second or third nucleotide immediately following the target site sequence complementary to LTG (i.e., there is a 1-2 nucleotide gap between where LTG and RTG bind). The 5' to 3' sequence of the LDG in the donor binding loop is complementary to the first strand of the donor site sequence. The 5' to 3' sequence of the RDG is the reverse complement of the first strand of the donor site sequence, and the 3' end of the RDG is the reverse complement of the nucleotide immediately following the target site sequence complementary to the LDG. In some embodiments, the 5' to 3' sequence of the RDG is the reverse complement of the first strand of the donor site sequence, and the 3' end of the RDG is the reverse complement of the second or third nucleotide immediately following the target site sequence complementary to the LDG (i.e., there is a 1-2 nucleotide gap between where the LDG and RDG bind). The first strand of the target site sequence and the first strand of the donor site sequence are on the same strand of the same DNA molecule.The donor and target sites can range from shorter DNA fragments (e.g., from about 11 bases or base pairs) up to multiple kilobases of DNA (e.g., the length of a chromosome) (e.g., lengths of about 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118 44, 45, 46, 47, 48, 49, 50, 75, or 100 b or bp to about 110, 120, 125, 150, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 525, 550, 575, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1250, 1500, 1750, 2000 , 2250, 2500, 2750, 3000, 3250, 3500, 3750, 4000, 4250, 4500, 4750, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10,000, 10,500, 11,000, 11,500, 12,000, 12,500, 13,000, 13,500, 14,000, 14,500, The excisions may be spaced any distance along the DNA molecule that results in excision of up to 15,000, 16,000, 17,000, 18,000, 19,000, 20,000, 21,000, 22,000, 23,000, 24,000, 25,000, 26,000, 27,000, 28,000, 29,000, 30,000 b or bp, or up to the length of the chromosome in the cell of interest).

[0230] In some embodiments where a core sequence is present, to mediate excision recombination, the 5' to 3' sequence of the LTG in the target binding loop is complementary to the first strand of the target site sequence, and the 3' end of the LTG is complementary to at least one nucleotide of the core sequence on the first strand of the target site sequence. The 5' to 3' sequence of the RTG is reverse complementary to the opposite strand of the target site sequence from the first strand, and the 3' end of the RTG is reverse complementary to at least one nucleotide of the core sequence on the opposite strand of the target site sequence from the first strand. In some embodiments, there is a 1-2 nucleotide gap between the binding sites of the LTG and RTG. The 5' to 3' sequence of the LDG is complementary to the first strand of the donor site sequence, and the 3' end of the LDG is complementary to at least one nucleotide of the core sequence on the opposite strand of the donor site sequence from the first strand. Th...

Claims

1. a) an IS110 family transposase or a nucleic acid comprising a sequence encoding said IS110 family transposase; and b) a nucleic acid comprising a sequence encoding a bridge RNA; A recombinant nucleic acid editing system comprising:

2. 2. The recombinant nucleic acid editing system of claim 1, wherein the nucleic acid comprising the sequence encoding the bridge RNA comprises the left end (LE) sequence of a transposon encoding the IS110 family transposase.

3. 2. The recombinant nucleic acid editing system of claim 1, wherein the nucleic acid comprising the sequence encoding the bridge RNA comprises the right end (RE) sequence of a transposon encoding the IS110 family transposase.

4. The recombinant nucleic acid editing system of claims 1 to 3, further comprising a nucleic acid comprising the RE sequence and LE sequence, or the RE sequence, core sequence, and LE sequence, of an IS110 element encoding the IS110 family transposase.

5. The recombinant nucleic acid editing system of claims 1 to 3, further comprising a nucleic acid comprising the right flanking (RF) sequence and left flanking (LF) sequence of the target site sequence of the IS110 family transposase, or the RF sequence, core sequence, and LF sequence.

6. The recombinant nucleic acid editing system of claim 4, wherein the nucleic acid comprising the RE sequence and the LE sequence, or the RE sequence, the core sequence, and the LE sequence, further comprises a nucleic acid sequence for insertion into a target site sequence.

7. The recombinant nucleic acid editing system of claim 6, wherein the target site sequence comprises an RF sequence and an LF sequence, or an RF sequence, a core sequence, and an LF sequence for the IS110 family transposase.

8. The recombinant nucleic acid editing system of claim 5, wherein the nucleic acid comprising the RF sequence and the LF sequence, or the RF sequence, the core sequence, and the LF sequence, further comprises a nucleic acid sequence for insertion into a donor site sequence.

9. The recombinant nucleic acid editing system of claim 8, wherein the donor site sequence comprises the RE sequence and LE sequence, or the RE sequence, core sequence, and LE sequence, of an IS110 element encoding the IS110 family transposase.

10. The recombinant nucleic acid editing system of claims 1 to 9, wherein the bridge RNA comprises a nucleotide sequence that is at least 50% identical to a bridge RNA sequence of SEQ ID NO: 1 to 348 or SEQ ID NO: 349 to 10175.

11. The RE sequence comprises an RE sequence of SEQ ID NO: 1 to 348, 30354 to 30529, 349 to 10175 or 30530 to 40356, the LE sequence comprises an LE sequence of SEQ ID NO: 1 to 348, 30354 to 30529, 349 to 10175 or 30530 to 40356, and / or the core sequence comprises a core sequence of SEQ ID NO: 1 to 348, 30354 to 30529, 349 to 10175 or 30530 to 40356. A recombinant nucleic acid editing system as described in claims 1 to 9.

12. a) an IS110 family transposase or a nucleic acid comprising a sequence encoding said IS110 family transposase; and b) a bridge RNA or a nucleic acid comprising a sequence encoding said bridge RNA; A recombinant nucleic acid editing system comprising: the bridge RNA comprises at least one stem-loop structure and further comprises at least one internal loop comprising a first nucleotide sequence complementary to a first target site sequence of a target DNA and a second nucleotide sequence complementary to a second target site sequence on an opposite strand of the target DNA to the first target site sequence, and the bridge RNA is capable of forming a complex with the IS110 family transposase. The recombinant nucleic acid editing system.

13. 13. The recombinant nucleic acid editing system of Claim 12, wherein the bridge RNA further comprises a third nucleotide sequence complementary to a first donor site sequence of the donor DNA and a fourth nucleotide sequence complementary to a second donor site sequence on the opposite strand of the donor DNA as the first donor site sequence.

14. The recombinant nucleic acid editing system of claim 13, wherein the third nucleotide sequence and the fourth nucleotide sequence are located on a second internal loop.

15. The recombinant nucleic acid editing system of claims 1 to 14, wherein the IS110 family transposase comprises a RuvC-like DEDD catalytic domain and a transposase domain.

16. 16. The recombinant nucleic acid editing system of claim 15, wherein the IS110 family transposase further comprises a linker domain between the RuvC-like DEDD catalytic domain and the transposase domain.

17. The recombinant nucleic acid editing system of claim 16, wherein the linker domain comprises a coiled-coil linker domain.

18. The recombinant nucleic acid editing system of claims 15 to 17, wherein the RuvC-like DEDD catalytic domain comprises an amino acid sequence at least 50% identical to a RuvC-like DEDD catalytic domain sequence of SEQ ID NOs: 10176-10523, 10524-20350, or 40357-516430.

19. The recombinant nucleic acid editing system of claims 15 to 17, wherein the IS110 family transposase comprises a RuvC-like DEDD catalytic domain that forms a tertiary structure similar to the RuvC-like DEDD catalytic domain of IS621.

20. The recombinant nucleic acid editing system of claim 19, wherein the IS110 family transposase RuvC-like DEDD catalytic domain comprises a tertiary structure similar to the tertiary structure of the IS621 RuvC-like DEDD catalytic domain when the template modeling score (TM score) of the IS110 family transposase RuvC-like DEDD catalytic domain is 0.5 or greater.

21. 21. The recombinant nucleic acid editing system of claim 20, wherein the RuvC-like DEDD catalytic domain comprises an amino acid sequence at least 15% identical to a RuvC-like DEDD catalytic domain sequence of SEQ ID NOs: 10176-10523, 10524-20350, or 40357-516430.

22. The recombinant nucleic acid editing system of claims 15 to 17, wherein the transposase domain comprises an amino acid sequence that is at least 50% identical to a transposase domain sequence of SEQ ID NOs: 10176 to 10523, 10524 to 20350, or 40357 to 516430.

23. The recombinant nucleic acid editing system of claims 15 to 17, wherein the IS110 family transposase comprises a transposase domain that forms a tertiary structure similar to the transposase domain of IS621.

24. The recombinant nucleic acid editing system of claim 23, wherein if the template modeling score (TM score) of the transposase domain of the IS110 family transposase is 0.5 or greater, the IS110 family transposase domain comprises a tertiary structure similar to the tertiary structure of the transposase domain of IS621.

25. 25. The recombinant nucleic acid editing system of claim 24, wherein the transposase domain comprises an amino acid sequence at least 15% identical to a transposase domain sequence of SEQ ID NOs: 10176-10523, 10524-20350, or 40357-516430.

26. The recombinant nucleic acid editing system of claims 15 to 17, wherein the IS110 family transposase comprises an amino acid sequence that is at least 50% identical to SEQ ID NOs: 10176-10523, 10524-20350, or 40357-516430.

27. The recombinant nucleic acid editing system of claims 15 to 17, wherein the IS110 family transposase further comprises a tertiary structure similar to the tertiary structure of IS621.

28. 28. The recombinant nucleic acid editing system of claim 27, wherein the IS110 family transposase comprises a tertiary structure similar to that of IS621 if the template modeling score (TM score) of the transposase is 0.5 or greater.

29. 29. The recombinant nucleic acid editing system of claim 28, wherein the transposase domain comprises an amino acid sequence at least 15% identical to a transposase domain sequence of SEQ ID NOs: 10176-10523, 10524-20350, or 40357-516430.

30. A recombinant nucleic acid editing system according to claims 1 to 29, wherein the IS110 family transposase is an IS110 group transposase.

31. A recombinant nucleic acid editing system according to claims 1 to 29, wherein the IS110 family transposase is an IS1111 group transposase.

32. The recombinant nucleic acid editing system of claim 31, wherein the IS1111 group transposase is IS1111A or IS1111_229727.

33. 31. The recombinant nucleic acid editing system of claim 30, wherein the IS110 group transposase is IS621, ISPa11, IsPa29, ISMmg1, ISPfl1, ISMae40, ISStma6, ISAzs32, ISMex9, ISCARN28, ISAar16, ISCps7, ISPpu9, ISRel9, ISEsa2, ISMma5, IS900, or ISHne5.

34. 31. The recombinant nucleic acid editing system of Claim 30, wherein the IS110 group transposase comprises an amino acid sequence that is at least 50% identical to IS621 (SEQ ID NO: 10176). 。

35. 35. The recombinant nucleic acid editing system of claims 12 to 34, wherein the bridge RNA comprises at least two stem-loop structures comprising a first stem-loop and a second stem-loop, the first stem-loop being 5' to the second stem-loop, the first stem-loop comprising a target-binding loop, and the second stem-loop comprising a donor-binding loop.

36. 36. The recombinant nucleic acid editing system of Claim 35, wherein the bridge RNA further comprises a third stem-loop structure 5' to the first stem-loop.

37. 37. The recombinant nucleic acid editing system of claims 35-36, wherein the stem of the first stem-loop is 5-35 nucleotides in length, the loop is 3-10 nucleotides in length, the target binding loop is 5-20 nucleotides in length, the stem of the second stem-loop is 5-35 nucleotides in length, the loop is 3-10 nucleotides in length, and the donor binding loop is 5-20 nucleotides in length.

38. 38. The recombinant nucleic acid editing system of Claim 37, wherein the stem of the second stem-loop structure comprises 1 to 4 loops or bubbles, each 1 to 10 nucleotides in length.

39. A recombinant nucleic acid editing system as described in claims 12 to 34, wherein the bridge RNA comprises a nucleotide sequence comprising any of the 5' to 3' sequences shown in Figure 19, where "n" represents any nucleotide, "R" represents an A or G nucleotide, and "Y" represents a C or U nucleotide.

40. The recombinant nucleic acid editing system of claim 39, wherein the bridge RNA comprises a 5' to 3' secondary structure shown in the first, second, third, or fourth column of the secondary structure for the sequence shown in Figure 19, where matching brackets "(" and ")" indicate base-paired nucleotides and "." indicates an unpaired base.

41. A recombinant nucleic acid editing system described in claims 12 to 34, wherein the bridge RNA comprises a stem-loop structure depicted in Figure 2D, Figure 11B, or Figure 13.

42. the target binding loop of the bridge RNA a left target guide (LTG) comprising, in a 5' to 3' direction, a nucleotide sequence complementary to a first strand of a target site sequence, wherein the 3' end of the LTG is complementary to at least one nucleotide of a core sequence on the first strand of the target site sequence; and a right target guide (RTG) comprising, in a 5' to 3' direction, a nucleotide sequence that is the reverse complement of the target site sequence on the opposite strand from the first strand, wherein the 3' end of the RTG is the reverse complement of at least one nucleotide of the core sequence on the opposite strand of the target site sequence from the first strand; Including, the target site sequence is a polynucleotide sequence; and / or the donor binding loop of the bridge RNA a left donor guide (LDG) comprising, in a 5' to 3' direction, a nucleotide sequence complementary to a first strand of a donor site sequence, wherein the 3' end of the LDG is complementary to at least one nucleotide of a core sequence on the first strand of the donor site sequence; and in the 5' to 3' direction to the opposite strand of the donor site sequence from the first strand. a right donor guide (RDG) comprising a nucleotide sequence that is a reverse complement of at least one of the nucleotides of the core sequence on the strand opposite the first strand of the donor site sequence, wherein the 3' end of the RDG is reverse complementary to at least one of the nucleotides of the core sequence on the strand opposite the first strand of the donor site sequence. Including, the donor site sequence is a polynucleotide sequence, and the core sequence of the donor site sequence is the same as the core sequence of the target site sequence; A recombinant nucleic acid editing system according to claims 12 to 41.

43. the target binding loop of the bridge RNA a left target guide (LTG) comprising, in a 5' to 3' direction, a nucleotide sequence that is reverse-complementary to the opposite strand of a target site sequence from the first strand, wherein the 3' end of the LTG is complementary to at least one nucleotide of a core sequence on the opposite strand of the target site sequence from the first strand; and a right target guide (RTG) comprising, in a 5' to 3' direction, a nucleotide sequence complementary to the first strand of the target site sequence, wherein the 3' end of the RTG is complementary to at least one nucleotide of the core sequence on the first strand of the target site sequence; Including, the target site sequence is a polynucleotide sequence; and / or the donor binding loop of the bridge RNA a left donor guide (LDG) comprising, in a 5' to 3' direction, a nucleotide sequence complementary to a first strand of a donor site sequence, wherein the 3' end of the LDG is complementary to at least one nucleotide of a core sequence on the first strand of the donor site sequence; and a right donor guide (RDG) comprising, in a 5' to 3' direction, a nucleotide sequence that is the reverse complement of the donor site sequence on the opposite strand to the first strand, wherein the 3' end of the RDG is reverse complementary to at least one nucleotide of the core sequence on the opposite strand of the donor site sequence to the first strand. Including, the donor site sequence is a polynucleotide sequence, and the core sequence of the donor site sequence is the same as the core sequence of the target site sequence; A recombinant nucleic acid editing system according to claims 12 to 41.

44. The target site sequence is sequence X 1 X 2 X 3 X 4 X 5 X 6 X 7 X 8 X 9 X 10 X 11 X 12 X 13 X 14 (wherein X is any nucleotide, and X 8 X 9 is the core, and X 12 , X 13 , and X 14 wherein one or more of the target site sequences are optionally part of the target site sequence.

45. The bridge RNA is arranged in the 5' to 3' direction as 1 X 2 X 3 X 4 X 5 X 6 X 7 X 8 LTG and Y in the 5' to 3' direction 14 Y 13 Y 12 Y 11 Y 10 Y 9 Y 8 and RTG, where Y is the complementary nucleotide to X, and Y 14 , Y 13 , and Y 12 one or more of which are optionally part of an RTG, or the bridge RNA is 1 X 2 X 3 X 4 X 5 X 6 X 7 X 8 X 9 LTG and Y in the 5' to 3' direction 14 Y 13 Y 12 Y 11 Y 10 Y 9 and RTG, where Y is the complementary nucleotide to X, and Y 14 , Y 13 , and Y 12 45. The recombinant nucleic acid editing system of claim 44, wherein one or more of:

46. The donor site sequence is the sequence STIR-X n1 -X 1 X 2 X 3 X 4 X 5 X 6 X 7 X 8 X 9 X 10 X 11 X 12 X 13 X 14 -X n2 -STIR (where X is any nucleotide, X 12 , X 13 , and X 14 one or more of the following are optionally part of the donor site sequence; the presence of STIR is optional, but if present, is a subterminal inverted repeat sequence containing 2 to 20 nucleotides; X 8 X 9 is a core, and n1 and n2 can independently be 0 to 10.

47. The bridge RNA is arranged in the 5' to 3' direction as 1 X 2 X 3 X 4 X 5 X 6 X 7 X 8 LDG and Y in the 5' to 3' direction 14 Y 13 Y 12 Y 11 Y 10 Y 9 Y 8 and RDG, where Y is the complementary nucleotide to X, and Y 14 , Y 13 , and Y 12 one or more of which are optionally part of an RDG, or the bridge RNA is 1 X 2 X 3 X 4 X 5 X 6 X 7 X 8 X 9 LDG and Y in the 5' to 3' direction 14 Y 13 Y 12 Y 11 Y 10 Y 9 and RDG, where Y is the complementary nucleotide to X, and Y 14 , Y 13 , and Y 12 47. The recombinant nucleic acid editing system of claim 46, wherein one or more of:

48. A recombinant nucleic acid editing system according to claims 12 to 47, wherein the target site sequence is located on genomic DNA, linear dsDNA, dsDNA plasmid, ssDNA or RNA.

49. A recombinant nucleic acid editing system according to claims 13 to 47, wherein the donor site sequence is located on genomic DNA, linear dsDNA, dsDNA plasmid, ssDNA or RNA.

50. The recombinant nucleic acid editing system of claims 48 to 49, wherein the dsDNA plasmid further comprises a polynucleotide sequence for insertion into the donor site sequence.

51. The recombinant nucleic acid editing system of claims 48 to 49, wherein the dsDNA plasmid further comprises a polynucleotide sequence for insertion into the target site sequence.

52. A recombinant nucleic acid editing system as described in claims 48 to 49, wherein the target site sequence and donor site sequence on the genomic DNA are located on the same DNA strand.

53. A recombinant nucleic acid editing system as described in claims 48 to 49, wherein the target site sequence and the donor site sequence on the genomic DNA are located on different chromosomes.

54. A recombinant nucleic acid editing system described in claims 12 to 53, wherein the bridge RNA is a split bridge RNA.

55. 55. The recombinant nucleic acid editing system of Claim 54, wherein the bridge RNA comprises a first RNA molecule comprising a first portion of the bridge RNA and a second RNA molecule comprising a second portion of the bridge RNA.

56. 56. The recombinant nucleic acid editing system of Claim 55, wherein the first portion of the bridge RNA comprises an internal loop comprising the first nucleotide sequence that is complementary to the first target site sequence of the target DNA, and the second portion of the bridge RNA comprises the second internal loop comprising the third nucleotide sequence that is complementary to the first donor site sequence of donor DNA and the fourth nucleotide sequence that is complementary to the second donor site sequence.

57. 57. The recombinant nucleic acid editing system of claims 55-56, wherein the first portion of the bridge RNA is encoded on a nucleic acid and is operably linked to a first promoter, and the second portion of the bridge RNA is encoded on the same or a different nucleic acid and is operably linked to a second promoter.

58. 57. The recombinant nucleic acid editing system of claims 55-56, wherein the first and second portions of the bridge RNA are encoded on a nucleic acid and further comprise one or more ribozyme sites such that the expressed bridge RNA is cleaved into the first and second RNA molecules.

59. A recombinant nucleic acid editing system described in claims 42 to 58, wherein any of the LTG, RTG, LDG, and / or RDG of the bridge RNA has a greater number of nucleotides complementary to their respective target site sequence or donor site sequence than the number of complementary nucleotides in its corresponding naturally occurring bridge RNA.

60. 60. The recombinant nucleic acid editing system of Claim 59, wherein the total number of nucleotides comprising the target binding loop and / or donor binding loop of the bridge RNA is the same as its corresponding naturally occurring bridge RNA.

61. 60. The recombinant nucleic acid editing system of Claim 59, wherein the total number of nucleotides comprising the target binding loop and / or donor binding loop of the bridge RNA is increased compared to its corresponding naturally occurring bridge RNA.

62. A recombinant nucleic acid editing system described in claims 42 to 58, wherein any of the LTG, RTG, LDG, and / or RDG of the bridge RNA is not complementary to nucleotides of the core sequence in their respective target site sequence or donor site sequence, but the number of complementary nucleotides in each target site sequence or donor site sequence is the same as in the corresponding naturally occurring bridge RNA.

63. The target site sequence is sequence X -1 X 1 X 2 X 3 X 4 X 5 X 6 X 7 X 8 X 9 X 10 X 11 X 12 X 13 X 14 (wherein X is any nucleotide, and X 8 X 9 is the core, and X 12 , X 13 , and X 14 wherein one or more of the X -1 X 1 X 2 X 3 X 4 X 5 X 6 X 7 LTG and Y in the 5' to 3' direction 14 Y 13 Y 12 Y 11 Y 10 Y 9 Y 8 and RTG, where Y is the complementary nucleotide to X, and Y 14 , Y 13 , and Y 12 44. The recombinant nucleic acid editing system of claim 42 or claim 43, wherein one or more of:

64. The donor site sequence is the sequence STIR-X n1 -X 1 X 2 X 3 X 4 X 5 X 6 X 7 X 8 X 9 X 10 X 11 X 12 X 13 X 14 -X n2 -STIR (where X is any nucleotide, X 12 , X 13 , and X 14 one or more of the following are optionally part of the donor site sequence; the presence of STIR is optional, but if present, is a subterminal inverted repeat sequence containing 2 to 20 nucleotides; X 8 X 9 is a core, and n1 and n2 can independently be 1 to 10), and the bridge RNA is -1 X 1 X 2 X 3 X 4 X 5 X 6 X 7 LDG and Y in the 5' to 3' direction 14 Y 13 Y 12 Y 11 Y 10 Y 9 Y 8 and RDG, where Y is the complementary nucleotide to X, and Y 14 , Y 13 , and Y 12 44. The recombinant nucleic acid editing system of claim 42 or claim 43, wherein one or more of:

65. the LTG, RTG, LDG, and / or RDG nucleotides of the bridge RNA 59. The recombinant nucleic acid editing system of claims 42 to 58, wherein one or more of: base pairs to nucleotides of their respective target site sequence or donor site sequence via non-canonical base pairing.

66. A vector comprising any one of the nucleic acid editing systems described in claims 1 to 65.

67. 67. A host cell comprising any of the vectors of claim 66.

68. A recombinant nucleic acid editing system described in claims 1 to 65, wherein any of the nucleic acids of the nucleic acid editing system further comprises an inducible promoter.

69. A method for integrating a target DNA molecule into a sequence-specific site of a target DNA in a cell, comprising introducing into the cell a nucleic acid editing system described in claims 1 to 65.

70. 70. The method of claim 69, wherein the cell is a mammalian cell.

71. 70. The method of claim 69, wherein the cell is a human cell.

72. 72. The method of claims 69 to 71, wherein the DNA of interest in the cell comprises a donor site sequence and the DNA molecule of interest is part of a nucleic acid of the nucleic acid editing system that comprises a target site sequence.

73. 72. The method of claims 69 to 71, wherein the DNA of interest of the cell comprises a target site sequence, and the DNA molecule of interest is part of a nucleic acid of the nucleic acid editing system that comprises a donor site sequence.

74. 73. The method of Claim 72, wherein the DNA of interest in the cell further comprises a second donor site sequence, the DNA molecule of interest further comprises a second target site sequence, and the nucleic acid editing system comprises a second bridge RNA that targets the second donor site sequence and the second target site sequence.

75. 74. The method of Claim 73, wherein the DNA of interest in the cell further comprises a second target site sequence, the DNA molecule of interest further comprises a second donor site sequence, and the nucleic acid editing system comprises a second bridge RNA that targets the second donor site sequence and the second target site sequence.

76. The method of any one of claims 72 to 75, wherein the sequence of the bridge RNA is engineered to bind to the donor site sequence and the target site sequence prior to introduction of the nucleic acid editing system.

77. The method according to any one of claims 70 to 76, wherein the target DNA of the cell is the genome of the cell.

78. The method of any one of claims 70 to 76, wherein the target DNA of the cell is a plasmid.

79. A method for inverting a DNA sequence of a target DNA in a cell, comprising introducing into the cell a nucleic acid editing system described in claims 1 to 65, wherein the target site sequence and the donor site sequence are present on the same target DNA molecule, and the LD of the donor site sequence and the RT of the target site sequence are on the same DNA strand.

80. 80. The method of claim 79, wherein the target DNA of the cell is the genome of the cell.

81. 81. The method of Claim 79 or Claim 80, wherein the sequence of the bridge RNA is modified to bind to the donor site sequence and the target site sequence prior to introduction of the nucleic acid editing system.

82. A method for excising a DNA sequence of a target DNA in a cell, comprising introducing into the cell a nucleic acid editing system described in claims 1 to 65, wherein a target site sequence and a donor site sequence are present on the same target DNA molecule, and the LD of the donor site sequence and the LT of the target site sequence are on the same DNA strand.

83. 83. The method of claim 82, wherein the target DNA of the cell is the genome of the cell.

84. The method of claims 82 to 83, wherein the sequence of the bridge RNA is modified to bind to the donor site sequence and the target site sequence prior to introduction of the nucleic acid editing system.

85. A method for translocating a DNA sequence between two linear DNA molecules of interest, comprising introducing into a cell a nucleic acid editing system described in claims 1 to 65, wherein a donor site sequence is present on a first linear DNA molecule and a target site sequence is present on a second linear DNA molecule.

86. 86. The method of claim 85, wherein the linear DNA molecule of interest in the cell is a chromosome of the cell.

87. The method of claims 85 to 86, wherein the sequence of the bridge RNA is modified to bind to the donor site sequence and the target site sequence prior to introduction of the nucleic acid editing system.