Novel RNA programmable system targeting polynucleotides
Patent Information
- Application Number
- JP2026081023
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-05-13
- Publication Date
- 2026-09-08
Smart Images

Figure 2026143439000001_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure relates to the field of targeted polynucleotide modification and detection of target sequences in polynucleotides. [Background technology]
[0002] Microorganisms are a fascinating and useful source of tools for genetic engineering, with applications across a wide range of technologies, including medicine.
[0003] In recent years, particular interest has been focused on adapting CRISPR-Cas systems derived from the adaptive immune systems of bacteria and archaea for use in DNA modification. Natural CRISPR-Cas systems are highly diverse and can be classified into two classes, each currently containing three distinct types with multiple subtypes (recently reviewed by Makarova et al., 2020). The most studied of these is the CRISPR-Cas9 (class 1, type II)-based system (reviewed by Jiang and Doudna, 2017), and more recently, the CRISPR-Cas12 (class 2, type V)-based system (described, e.g., Zetsche et al., 2015 and Karvelis et al., 2020). These systems are based on RNA-guided DNA endonucleases, configured as ribonucleoprotein (RNP) complexes capable of introducing site-specific double-strand breaks into target DNA. In both cases, target recognition by the RNP complex requires the presence of a short protospacer-adjacent motif (PAM) that flanks at the target site. Nevertheless, these systems are very attractive as a source of tools for DNA modification due to their ability to induce endonucleases to different targets by altering the RNA sequence (when a suitable PAM is present). [Overview of the initiative]
[0004] Surprisingly, this invention demonstrates that the TnpB protein is an RNA-binding protein that forms a ribonucleoprotein effector complex with an RNA molecule. Furthermore, it was shown that these effector complexes are capable of cleaving polynucleotides based on the binding of a segment of the RNA molecule to a target sequence in the polynucleotide, and the subsequent nuclease activity of the TnpB protein in the complex.
[0005] The TnpB protein is known in the art as the predicted product of the tnpB gene found in several families of bacterial and archaeal insertion sequences (IS). Insertions are widely found prokaryotic mobile genes, and many eukaryotic DNA transposition factors are associated with them (Hickman et al., 2010). Insertions contain only genes related to transposition and its regulation. Several families of insertion sequences hold both tnpA and tnpB genes. However, while the function of TnpA in transposition is well established, the role of TnpB has not been demonstrated. TnpB has been shown to be non-essential for transposition, and the protein is thought to be involved in the negative regulation of transposon excision and insertion (Kersulyte et al., 2000, 2002; Pasternak et al., 2013). Previously, it had not been shown that the TnpB protein can act as a nuclease when bound to RNA, nor that cleavage is targeted by binding between RNA segments and target sites.
[0006] The experiments described herein demonstrate that the TnpB protein can be used to produce a novel RNA-guided effector complex that is functionally distinct from the conventional CRISPR-Cas9 and CRISPR-Cas12 systems, in which the TnpB protein can act as a nuclease. Unlike the CRISPR-Cas systems, there is no CRISPR array to associate with the insertion sequence. Rather, surprisingly, it was shown that the RNA to which the TnpB protein associates in nature originates from the portion of the insertion sequence.
[0007] The effector complexes described herein have significant utility in polynucleotide targeting in vitro, ex vivo, or in vivo, and advantageously expand the genetic modification toolbox. Not only can polynucleotides be modified by utilizing the nuclease activity of the TnpB protein, but the TnpB protein can also be mutated to inactivate its nuclease activity, allowing the effector complex to be used to block gene expression or to detect target sequences without polynucleotide cleavage.
[0008] Furthermore, the effector complexes described herein, including active or inactive forms of the TnpB protein, can be designed to hold one or more additional effector molecules to a target site within a polynucleotide. In some examples, the TnpB protein or its inactive form may be contained in a fusion protein with one or more effector molecules.
[0009] Because TnpB proteins are relatively small in size, they are particularly suitable for use in cell delivery via AAV-based delivery and in therapeutic applications, for example. In certain situations where the size of the effector complex is important, the TnpB-based effector complex of the present invention is advantageous compared to the larger Cas9 and Cas12 proteins, which are 1000–1500 amino acid lengths and 500–1500 amino acid lengths, respectively.
[0010] Therefore, the present invention provides the following:
[0011] In a first embodiment, the present invention provides a method for cleaving a polynucleotide using an effector complex, wherein the polynucleotide comprises a target sequence, and the effector complex comprises: (a) Proteins containing or consisting of TnpB protein; and (b) RNA containing the following: (i) a polynucleotide targeting segment containing a guide sequence capable of hybridizing to a target sequence; and (ii) A protein-binding segment that enables RNA to bind to the TnpB protein to form an effector complex, wherein the method comprises contacting a polynucleotide with the effector complex and enabling the TnpB protein to cleave the polynucleotide.
[0012] In a second embodiment, the present invention provides RNA for guiding an effector complex to a target region in a polynucleotide, the RNA comprising: (i) a polynucleotide targeting segment comprising a guide sequence capable of hybridizing to a target sequence in the target region of a polynucleotide; and (ii) A protein-binding segment that enables RNA to bind to the TnpB protein and form an effector complex.
[0013] In a third aspect, the present invention provides an effector complex for binding to a target region in a polynucleotide, the effector complex comprising a protein and RNA, wherein the protein comprises or consists of the TnpB protein, and the RNA comprises: (i) a polynucleotide targeting segment containing a guide sequence capable of hybridizing with a target sequence contained in the target region; and (ii) Protein-binding segment that binds to the TnpB protein.
[0014] In a fourth embodiment, the present invention provides a fusion protein comprising a TnpB protein, and (i) one or more nuclear localization signal and / or cell membrane permeable peptides on the amino or carboxyl terminus of the fusion protein, and / or (ii) one or more effector molecules.
[0015] In a fifth aspect, the present invention provides a mutant TnpB protein optionally comprising a mutation for inactivating the nuclease domain of the protein, wherein the mutant TnpB protein is the TnpB protein of the fusion protein of the present invention.
[0016] In a sixth aspect, the present invention provides a DNA encoding an RNA.
[0017] In a seventh aspect, the present invention provides a DNA or RNA encoding a fusion protein.
[0018] In an eighth aspect, the present invention provides a DNA or RNA encoding a mutant TnpB protein.
[0019] In a ninth aspect, the present invention provides a recombinant expression vector comprising the DNA of the present invention.
[0020] In a tenth aspect, the present invention provides a host cell comprising the recombinant expression vector of the present invention or the DNA of the present invention.
[0021] In an eleventh aspect, the present invention provides a composition comprising the RNA of the present invention, the effector complex of the present invention, the fusion protein of the present invention, the mutant TnpB protein of the present invention, the DNA of the present invention, the recombinant expression vector of the present invention, or the host cell of the present invention, and a buffer.
[0022] In a twelfth aspect, the present invention provides an in vivo, ex vivo, or in vitro method for producing the RNA, effector complex, fusion protein, or mutant TnpB protein of the present invention.
[0023] In a thirteenth aspect, the present invention provides a system for modifying a target region in a polynucleotide, wherein the target region comprises a target sequence, and the system comprises: a) a protein comprising or consisting of a TnpB protein, or a DNA encoding said protein, and b) RNA or DNA encoding RNA, wherein the RNA comprises: (i) a polynucleotide targeting segment comprising a sequence complementary to a target sequence; and (ii) a protein binding segment that binds to a TnpB protein.
[0024] In a fourteenth aspect, the present invention provides an RNA, an effector complex, a mutant TnpB, a fusion protein, a DNA encoding any of the foregoing, or a system for use as a medicament or for use in a diagnostic method.
[0025] In a fifteenth aspect, the present invention provides use of an RNA, an effector complex, a mutant TnpB, a fusion protein, a DNA encoding any of the foregoing, or a system in an ex vivo or in vitro method for determining the presence of a polynucleotide comprising a target sequence in a sample.
[0026] In a sixteenth aspect, the present invention provides use of an RNA, an effector complex, a mutant TnpB, a fusion protein, a DNA encoding any of the foregoing, or a system in an in vivo, ex vivo, or in vitro method for modifying a target region of a polynucleotide, wherein the target region comprises the target sequence.
[0027] In a seventeenth aspect, the present invention provides use of an RNA, an effector complex, a mutant TnpB, a fusion protein, a DNA encoding any of the foregoing, or a system in an in vivo, ex vivo, or in vitro method for genetically modifying a cell.
[0028] In an eighteenth aspect, the present invention provides a genetically modified cell for use as a medicament in a subject, wherein the cell is obtained by a method comprising the step of genetically modifying a cell obtained from the subject using the system or effector complex of the present invention.
[0029] In a 19th aspect, the present invention provides a method for modifying, labeling, or controlling the expression of a target region in a polynucleotide using an effector complex, wherein the target region comprises a target sequence, wherein the effector complex is (i) the effector complex of the present invention; (ii) comprising a fusion protein and RNA of the present invention; or (iii) comprising a mutant TnpB protein, or a fusion protein comprising a mutant TnpB protein, and RNA of the present invention, wherein the method comprises the step of contacting a polynucleotide with the effector complex, thereby allowing the guide sequence of the RNA to hybridize to the target sequence, enabling the effector complex to modify or label the target region or control the expression from the target region. [Brief explanation of the drawing]
[0030] To aid in understanding this disclosure and to illustrate how embodiments may be carried out, the accompanying drawings are referenced merely as examples.
[0031] [Figure 1A] Figure 1 relates to the IS200 / IS605 mobile gene characterization. Figure 1A shows a schematic of the Deinococcus radiodurans ISDra2 locus. The system consists of tnpA and tnpB genes that flank to left and right partial palindromic sequences (LE and RE), respectively. [Figure 1B] This diagram outlines the transposon excision and paste mechanism mediated by TnpA. During replication, the TnpA dimer mediates transposon excision from the host DNA lagging strand, forming a circular single-stranded DNA intermediate and a donor joint. The excised transposon is then inserted into the next host lagging DNA strand of the short 5'-TTGAT-3'(ISDra2) motif at the acceptor joint, completing the transposition cycle. Transposon excision / insertion sites are indicated by triangles. [Figure 1C] This document outlines the experimental workflow for TnpB complex expression and its purification and binding RNA extraction from E. coli cells. [Figure 1D]This provides alignment of sRNA sequencing reads to the ISDra2 locus. Transposon excision / insertion sites are indicated by triangles. The RNA sequence derived from the RE factor consists of ribonucleotides that may be involved in hairpin formation, and two ribonucleotides between the hairpin and the triangle, while the last approximately 16nt of the 3' end of the sequenced RNA aligns with the transposon flanking DNA-ribonucleotide shown to the right of the triangle.
[0032] [Figure 2A] Figure 2 shows TnpB from ISDra2 system purification. Figure 2A shows an SDS-PAGE gel displaying the elution fraction of the protein bound to a HisTrap chelate column prepared from single TnpB expression and purification from E. coli cells. The framed area represents a band of expected 10×MBP-TnpB protein size (95.4 kDa). [Figure 2B] This SDS-PAGE gel shows the elution fraction of a protein bound to a HisTrap chelate column prepared from TnpB using ISDra2 system expression and purification from E. coli cells. The area enclosed by the frame represents a band of expected 10×MBP-TnpB protein size (95.4 kDa). [Figure 2C] This shows an SDS-PAGE gel of a pooled fraction containing 10×MBP-TnpB protein. [Figure 2D] This shows a gel related to the detection and analysis of nucleic acids co-purified with TnpB protein.
[0033] [Figure 3A]Figure 3 shows that the TnpB protein is an RNA-guided dsDNA nuclease. Figure 3A provides a schematic of the experimental workflow for detecting double-stranded (ds)DNA cleavage activity. The reRNA coding construct contains a 16nt guide sequence. F-Forward primers anneal to a ligated adapter. R1 and R2-Reverse primers anneal to the plasmid backbone. 7N represents the randomization region in the plasmid library following the target sequence. [Figure 3B] This shows the determination of the adapter ligation site that indicates double-strand break (DSB) formation in the target sequence. [Figure 3C] This provides WebLogo representations of motifs identified in 7N randomized regions in 20-21 bp F+R1 enriched adapter concatenated reads. [Figure 3D] This document provides an outline of the experimental workflow for TnpB RNP complex expression and purification. The reRNA coding construct includes a 16nt guide sequence. [Figure 3E] The gel provides evidence that the TnpB RNP complex cleaves the supercoiled target plasmid in vitro, unwinding the plasmid, and that the cleavage depends on an intact RuvC-like active site. [Figure 3F] This provides a gel demonstrating that targets complementary to the transposase association motif (TAM) and the reRNA 3' terminal sequence are required for plasmid DNA cleavage. [Figure 3G] This shows Sanger sequencing of TnpB cleavage plasmid products, revealing cleavage sites from 5'-TAM to 15-21 bp. The identified cleavage sites are marked using triangles (NTS - non-target strand; TS - target strand).
[0034] [Figure 4A]Figure 4 shows that the TnpB RNP complex cleaves dsDNA in a TAM-dependent manner. Figure 4A shows a schematic of the experimental workflow for detecting double-stranded (ds)DNA cleavage activity. The reRNA coding construct contains a 20nt guide sequence. Forward primers anneal to the F-linked adapter. Reverse primers anneal to the R1 and R2 plasmid backbones. 7N represents the randomization region in the plasmid library following the target sequence. [Figure 4B] This shows the determination of the adapter ligation site that indicates double-strand break (DSB) formation in the target sequence. [Figure 4C] This shows the WebLogo representation of motifs identified in the 7N randomized region of 20-21 bp F+R1 enriched adapter concatenated reads. [Figure 4D] This shows the WebLogo representation of motifs identified in the 7N randomized region of 20-21bp F+R1(-TnpB) enriched adapter concatenated reads.
[0035] [Figure 5A] Figure 5 shows that TnpB mediates plasmid interference in vivo. Figure 5A shows a schematic of the experimental workflow for a plasmid interference assay in E. coli. Cleavage of the target plasmid results in loss of resistance to kanamycin (Kn). The reRNA coding construct contains a 16nt guide sequence. AmpR - ampicillin / carbenicillin (Ap / Cb) resistance gene, KanR - Kn resistance gene. [Figure 5B] The results of the transformation experiment are shown. The transformation experiment involved serial dilution (10×), and the E. coli transformants were grown for 44 hours at 25°C on a medium supplemented with Cb and Kn.
[0036] [Figure 6A] Figure 6 shows the purification of the TnpB RNP complex. Figure 6A shows a schematic of the experimental workflow for TnpB RNP complex expression and multi-step purification. The reRNA coding construct includes a 16nt guide sequence. [Figure 6B]The results of SDS-PAGE analysis of purified TnpB and TnpB(D191A)RNP complexes are shown. [Figure 6C] This shows the molecular weights of the TnpB and reRNA RNP complex determined by mass photometry. The obtained molecular weights correspond to the TnpB RNP complex, which consists of a TnpB protein bound to approximately 150 nt of reRNA (1:1 molar ratio).
[0037] [Figure 7A] Figure 7 shows that TnpB nuclease is a novel genome editor. Figure 7A shows an overview of the experimental workflow for a human cell line (HEK293T) genome editing experiment. [Figure 7B] This shows the detection of indel activity at five tested 20 bp long targets in human genomic DNA (expressed as the mean of three replicates, ± standard deviation). Along the x-axis, for each site, the bar representing "TnpB (non-target)" is on the left-hand side, and the bar representing "TnpB" is on the right-hand side. [Figure 7C] The results of indel profile analysis at the EMX1-1 site, which shows the dominant deletion across the transection site, are shown. The shaded elongated area on the left of the graph represents the "TAM," and the shaded elongated area on the right represents the "target."
[0038] [Figure 8A] Figure 8 shows synthetic dsDNA cleavage by the TnpB RNP complex. Figure 8A provides a gel showing that the purified TnpB RNP complex cleaves a dsDNA substrate containing a target (shown in green) represented by the sequence CTCAGGGAACCGCGGG (SEQ ID NO: 17) (3'→5') on the TS (target strand) and a TAM (shown in red) represented by the sequence TTGAT (5'→3') on the NTS (non-target strand), generating a stepped cleavage pattern. NTS and TS represent the non-target and target strands, respectively. The TnpB(D191A)RNP complex is incubated with the D-DNA substrate for 60 minutes. [Figure 8B]This provides a gel demonstrating that the purified TnpB RNP complex does not cleave target-containing dsDNA substrates in the absence of double-stranded TAMs. The TnpB(D191A)RNP complex is incubated with a D-DNA substrate for 60 minutes.
[0039] [Figure 9A] Figure 9 shows synthetic ssDNA cleaved by the TnpB RNP complex. Figure 9A is a gel showing that the purified TnpB RNP complex cleaves an ssDNA substrate containing a sequence complementary to the reRNA target sequence. NTS and TS represent the non-target and target strands, respectively. The TnpB(D191A)RNP complex is incubated with a D-DNA substrate for 60 minutes. [Figure 9B] This gel demonstrates that purified TnpB RNP complexes cleave ssDNA substrates containing sequences complementary to the reRNA target sequence. NTS and TS represent the non-target and target strands, respectively. The TnpB(D191A)RNP complex is incubated with a D-DNA substrate for 60 minutes.
[0040] [Figure 10A] Figure 10 shows the results of in vitro TnpB cleavage condition tests. Figure 10A shows the results of assays determining TnpB RNP plasmid DNA cleavage at various temperatures. Products were analyzed after incubation of plasmid DNA with the TnpB RNP complex for 15 minutes. [Figure 10B] The results of an assay to determine TnpB RNP plasmid DNA cleavage at various NaCl concentrations are shown. Products were analyzed after incubation of plasmid DNA with the TnpB RNP complex for 15 minutes.
[0041] [Figure 11A]Figure 11 shows that TnpB mediates plasmid interference in vivo. Figure 11A provides a schematic of the experimental workflow for a plasmid interference assay in E. coli. Cleavage of the target plasmid results in loss of resistance to kanamycin (Kn). The reRNA coding construct contains a 16nt guide sequence. AmpR is the ampicillin / carbenicillin (Ap / Cb) resistance gene, and KanR is the Kn resistance gene. [Figure 11B] The transformation experiment was performed with serial dilutions (10×), and the results show that E. coli transformants grew on a medium supplemented with Cb and Kn at 25-37°C.
[0042] [Figure 12]This provides alignments of RuvC I, RuvC II, and RuvC III motifs of TnpB proteins from different insertion sequences. The motif sequences are: ISDra2 (IS605 family) TnpB protein (SEQ ID NO: 1); ISHp608 (IS605 family) TnpB protein (SEQ ID NO: 2); IS605 (IS605 family) TnpB protein (SEQ ID NO: 3); IS606 (IS605 family) TnpB protein (SEQ ID NO: 4); IS609 (IS605 family) TnpB protein (SEQ ID NO: 5); IS1341 (IS1341 family) TnpB protein (SEQ ID NO: 6); ISC1316 (IS1341 family) TnpB protein (SEQ ID NO: 7); IS891 (IS1341 family) TnpB protein (SEQ ID NO: 8); ISEc42 (IS1341 family) The TnpB proteins are obtained from Millie TnpB protein (SEQ ID NO: 9), ISTel3 (IS1341 family) TnpB protein (SEQ ID NO: 10), IS607 (IS607 family) TnpB protein (SEQ ID NO: 11), ISTsi1 (IS607 family) TnpB protein (SEQ ID NO: 12), IS1535 (IS607 family) TnpB protein (SEQ ID NO: 13), ISBlo12 (IS607 family) TnpB protein (SEQ ID NO: 14), and ISC1926 (IS607 family) TnpB protein (SEQ ID NO: 15). The alignment shows that the active site residues (D---E---D enclosed in a box) are conserved across the TnpB family.
[0043] [Figure 13] This disclosure outlines the nuclease activity of the RNA-guided ribonucleoprotein complex. The ribonucleoprotein complex (containing the TnpB protein and RNA) recognizes a double-stranded TAM sequence (located at 5' of the target sequence with reference to the non-target strand), and the RNA guide sequence in the ribonucleoprotein complex binds to the target sequence of the polynucleotide target strand (TS), thereby resulting in cleavage of the target strand (TS) and non-target strand (NTS) by the RuvC-like domain of the TnpB protein. [Modes for carrying out the invention]
[0044] The inventors have identified novel RNA-guided ribonucleoproteins (also referred to herein as “effector complexes”) that function in a manner similar to, but distinct from, Cas9 and Cas12 DNA endonucleases. Accordingly, this disclosure relates in particular to these effector complexes, methods involving their use in vitro, ex vivo, and in vivo (prokaryotic and eukaryotic cells) to cleave or modify polynucleotides, and systems for their delivery to target cells.
[0045] TnpB protein The proteins of this disclosure are proteins containing, essentially consisting of, or comprising the TnpB protein. In particular, where a protein "contains" the TnpB protein, further amino acids may be present in the protein. This includes fusion proteins of TnpB with one or more additional effector proteins, as further described below. Where a protein "essentially consists of" the TnpB protein, further amino acids or protein sequences may be present in the protein that do not significantly affect the essential features of the TnpB protein, namely its ability to bind to RNA and form the effector complex described herein (which may have the ability to act as an RNA programmable nuclease, where the TnpB protein retains its nuclease activity, or may have the ability to act as an RNA programmable carrier or RNA programmable polynucleotide blocker, where TnpB in the effector complex is an inactive / mutant TnpB protein in which its nuclease activity is inactivated, as further described below). Where a protein "consists of" the TnpB protein, further amino acids are not present.
[0046] TnpB proteins are proteins encoded by the tnpB gene from an insertion sequence (IS), or sequence variants of these TnpB proteins that retain the ability to form the effector complex described herein. In particular, in the examples of this disclosure, TnpB proteins have the amino acid sequence of a protein obtained from the tnpB gene of a mobile gene in the IS200 / IS605 or IS607 family, or sequence variants thereof. In preferred examples, TnpB proteins have the amino acid sequence of a protein obtained from the tnpB gene of a mobile gene in the IS200 / IS605 family, or sequence variants thereof. More specifically, TnpB proteins may have the amino acid sequence of a protein obtained from the tnpB gene of a mobile gene in the IS200 / IS605 family found in the bacterial Deinococcus family, or sequence variants thereof. In one example, the TnpB protein has the amino acid sequence of the TnpB protein obtained from the tnpB gene of ISDra2 (insertion sequence IS200 / IS605 from Deinococcus radiodurans) or a variant thereof.
[0047] As described above, insert sequences are simple, widely found mobile genetic elements (MGEs) containing only genes related to transposition and the regulation of transposition. Insert sequences are classified into different families in the art, as described in Siguier et al., 2006 and Siguier et al., 2014, and as shown in ISfinder (https: / / isfinder.biotoul.fr / ), a database providing lists of insert sequences isolated from bacteria and archaea. While the sequences of these insert sequences can be diverse, transposition elements of the IS200 / IS605 family are identified as those that hold terminal palindromic elements (LE and RE) at the ends of the MGE, and tnpA and tnpB genes in different configurations, or standalone tnpA or tnpB genes. In particular, the IS200 / IS605 family can be further classified into IS200 (retaining only the tnpA gene), IS200 / IS605 (sometimes also referred to as IS605, and retaining both the tnpA and tnpB genes, e.g., IS608 from Helicobacter pylori and ISDra2 from Deinococcus radiodurans (the composition of this factor is shown in Figure 1D)), and IS1341 (retaining only the tnpB gene). IS607 MGE is identified as encoding both the tnpA and tnpB genes, and its coding sequence is sometimes duplicated. The ends of these factors can also associate with inverted repeat sequences, which are often incomplete and / or secondary RNA structures.
[0048] The TnpB protein comprises an RNA-binding segment and a RuvC-like nuclease domain, which together enable the TnpB protein to form the effector complex described herein, the effector complex having nuclease activity against a target region (including a target site to which an RNA guide sequence binds) in a polynucleotide. In particular, as demonstrated herein, the RuvC-like domain is responsible for the nuclease activity of the TnpB protein.
[0049] RuvC itself is a dimeric bacterial endonuclease that dissociates Holliday junctions in bacteria and requires a divalent metal ion for activity. RuvC-like domains (containing RuvC-I, RuvC-II, and RuvC-III motifs, optionally accompanied by a Zn finger between the RuvC-II and RuvC-III motifs) are known in the art and are recognized as being responsible for single-strand cleavage by the Cas9 protein and double-strand nuclease activity by the Cas12 protein (see, e.g., Shmakov et al., 2017, Makarova et al., 2015, and Makarova et al., 2020). Similar to the RuvC protein, the RuvC-like domain of TnpB typically requires a divalent metal ion for activity.
[0050] Figure 12 provides alignments of the RuvC-I, RuvC-II, and RuvC-III motifs of the TnpB protein from different insert sequences. (The names and families of the insert sequences are shown on the left-hand side.) The alignments show the conserved D---E---D amino acids (the amino acids enclosed in boxes in Figure 12) in motifs I, II, and III, respectively, that are involved in the RuvC active site. These amino acids can be identified within the TnpB protein using a sequence alignment tool, such as the Clustal Omega sequence alignment program (https: / / www.ebi.ac.uk / Tools / msa / clustalo / ) (Madeira et al., 2019).
[0051] The polynucleotide containing the target sequence for which the TnpB protein has nuclease activity can be double-stranded DNA or single-stranded DNA. In particular, the TnpB protein has nuclease activity against double-stranded DNA, and therefore, effector complexes containing the TnpB protein are especially useful in genome editing.
[0052] The RNA-binding segment of the TnpB protein contains a sequence that interacts with RNA to form an effector complex. As shown in the experiments reported herein, the inventors found that expression of the tnpB gene fused to a maltose-binding protein sequence alone in E. coli, followed by affinity chromatography, revealed a low yield of intact TnpB protein. However, co-expression with RNA resulted in a higher yield of TnpB protein. While not wishing to be constrained by theory, the inventors believe that the interaction of the RNA-binding segment of the TnpB protein with RNA acts to stabilize the TnpB protein.
[0053] For TnpB to be able to cleave double-stranded DNA, the polynucleotide must contain a TnpB association sequence motif 5' of the target sequence (on the non-target strand as shown in Figure 13). This TnpB association sequence motif is also referred to herein as the transposon-associated motif or TAM. In particular, although we do not wish to be constrained by theory, the inventors assume that, in a manner similar to the requirements for Cas9 and Cas12 effector proteins for PAM, effector complex cleavage of a DNA molecule requires the presence of a TnpB association sequence motif, and that the TAM is recognized as a double-stranded motif by the effector complex (since its presence is not required for cleavage of single-stranded DNA by the effector complex, as exemplified herein). The sequence of the TAM is expected to vary among different TnpB proteins. The TAM sequence for a specific TnpB protein can be determined using a PAM (protospacer adjacent motif) recognition assay previously developed for Cas9 and Cas12 nucleases (Karvelis et al., 2015, 2019) (see also Example 2).
[0054] The TnB association sequence motif in a polynucleotide can be a T-rich motif and can be a TTGAT. Particularly preferably, the TnpB association sequence is a TTGAT, and the TnpB protein is derived from the ISDra2 family, and more preferably includes or consists of the amino acid sequence of SEQ ID NO: 1 or a sequence variant thereof.
[0055] The TnpB protein may be, i.e., derived from, the tnpB gene or its sequence variant found in the inserted sequence. The TnpB sequence variant retains an RNA-binding segment and a RuvC-like nuclease domain, which together enable the TnpB protein to form the effector complex described herein, having nuclease activity against a target region in a polynucleotide (including the target site where RNA binds). If the effector complex targets a polynucleotide that is double-stranded DNA, the TnpB protein variant must also retain the ability to recognize the TnpB association motif in the target region of the polynucleotide.
[0056] The sequence variant may have at least 85%, at least 90%, at least 95%, at least 98%, and at least 99% sequence identity with the TnpB protein produced from the tnpB gene from the IS family shown above. Alternatively, the variant may have at least 85%, at least 90%, at least 95%, at least 98%, and at least 99% sequence similarity with the TnpB protein (as determined in particular by BLAST).
[0057] Sequence variations can be based on the establishment of conserved amino acid changes. Furthermore, the methods described in the Art, which have been used to increase the specificity and activity of Cas9 and Cas12 proteins, can also be used, in particular, to generate TnpB variants with reduced off-target nuclease activity. One example is directed evolution.
[0058] TnpB proteins can be 300–600 amino acids long, optionally 350–550 amino acids long, and even more optionally 350–450 amino acids long.
[0059] For example, the TnpB protein may be a TnpB protein from the tnpB gene of ISDra2 (insertion sequence IS200 / IS605 from Deinococcus radiodurans), which is a 408 amino acid sequence having the amino acid sequence of SEQ ID NO: 1 (see https: / / isfinder.biotoul.fr / scripts / ficheIS.php?name=ISDra2 and NCBI accession number AE000513), or a TnpB protein with an amino acid sequence having at least 85% sequence identity with SEQ ID NO: 1, or at least 90%, at least 95%, at least 98%, or at least 99% sequence identity with it. MIRNKAFVVRLYPNAAQTELINRTLGSARFVYNHFLARRIAAYKESGKGLTYGQTSSELTLLKQAEETSWLSEVDKFALQNSLKNLETAYKNFFRTVKQSGKKV GFPRFRKKRTGESYRTQFTNNNIQIGEGRLKLPKLGWVKTKGQQDIQGKILNVTVRRIHEGHYEASVLCEVEIPYLPAAPKFAAGVDVGIKDFAIVTDGVRFKH EQNPKYYRSTLKRLRKAQQTLSRRKKGSARYGKAKTKLARIHKRIVNKRQDFLHKLTTSLVREYEIIGTEHLKPDNMRKNRRLALSISDAGWGEFIRQLEYKAAWYGRLVSKVSPYFPSSQLCHDCGFKNPEVKNLAVRTWTCPNCGETHDRDENAALNIRREALVAAGISDTLNAHGGYVRPASAGNGLRSENHATLVV(Sequence ID:1)
[0060] In further examples, the TnpB protein may be one of the following, or a sequence variant having at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% sequence identity with it.
[0061] ISHp608 (IS605 family) TnpB protein. Length: 383; NCBI accession number AF357224 IS Finder Database Entry: https: / / isfinder.biotoul.fr / scripts / ficheIS.php?name=ISHp608 MLITYKQKLYKNDKNRRIDTLLRRYGALYNHCIALHKRYYRLFKKYLKLYDLQKHITKLKKTHRYAFLKTLGSQTMQDLTERIDKAFKKFFNKKAKLPRFKKVANYKSFTFKSKIDKKTGLNKGVGFAIKDNVVSFNGYSYKFIKTYAFIGKVKTLTIKRDNTGDYFLCLVCELENHPNKQTACDKSVGFDFGLKTFLTGSDHTKIESPLSFSKYLPLIKRLSKNLSKKVKGSNNFKKAKKKLTQLHQKIKYLRTDFFHKLALKLSREYQTIFIEDLNMKAMQKLWGRKVSDLAFSEFVKILENKANVVKIDRFYPSSKTCSNCLFVNEEINKDFRKIGKTDKEREYHCKYCGLELDRDLNAAINIHRVGASTLGVEFVRPTC(Sequence ID:2)
[0062] IS605 (IS605 family) TnpB protein. Length: 427; from NCBI accession number HPU60177 IS Finder database entry: https: / / isfinder.biotoul.fr / scripts / ficheIS.php?name=IS605 MLNAIKFRIYPNAQQKELISKHFGCSRVVYNYFLDYRQKQYAKGIKETYFTMQKVLTQIKHQEKYHYLNECNSQSLQMALRQLVSAYDNFFSKRARYPKFKSKKNAKQ SFAIPQNIEIKTETQTIALPKFKEGIKAKLHRELPKDSVIKQAFISCIADQYFCSISYETKEPIPKPTIIKKAVGLDMGLRTLIVTSDKIEYPHIRFYQKLEKKLTKAQ RRLSKKVKGSNNRKKQAKKVARLHLACSNTRDDYLHKISNEITNQYDLIGVETLNVKGLMRTYHSKSLANASWGKFLTMLKYKAQRKAKTLLGIDRFFPSSQLCSYCGFNTGKKHENITKFTCPHCNITHHRDYNASVNIRNYALGMLDDRHKIKIDKSRVGIIRTDYAHYTDERIKACGASSNGVISKYGNILDLASYGAMKQEKAQSL(Sequence ID:3)
[0063] IS606 (IS605 family) TnpB protein. Length: 442. From NCBI accession number U95957. IS Finder Database Entry: https: / / isfinder.biotoul.fr / scripts / ficheIS.php?name=IS606 MKVNKGFKFRLYPTKEQQDKLQRCFFVYNQAYNIGLNLLQEQYETNKDSPPKERKWKKSSELDKAIKHHLNARGLSFSSVIAQQSRMNVERALKDAFKVKDRGFPKFKNSKS AKQSFSWNNQGFSIKDSDEERFKIFTLMKMPLMMRMHRDFPPHSKVKQIVISWSHRKYFVSFCVEYEQDITPIKNPKNGVGLDLNILDIACSCGVNNHKKLTDFKQYPTDMKE LLGIEIDEELDTKRLIPTYSKLYSLKKYSKKFKRLQRKQSRRVLKSKQNKTKLGGNFYKTQKKLNQAFDKSSHQKTDRYHKITSELSKQFELVVVEDLQVKNMTKRAKLKNVKQKSGLNQSILNTSFYQIISFLDYKQQHNGKLLVKVPPQYTSKTCHCCGNINHKLKLNHRQYWCLECGYREHRDINAANNIISKGLSLFGVGNIHADFKEQSLSC(Sequence ID: 4)
[0064] IS609 (IS605 family) TnpB protein. Length: 402; NCBI accession number BA000007 IS Finder Database Entry: https: / / isfinder.biotoul.fr / scripts / ficheIS.php?name=IS609 MKRLQAFKFQLRPGGQQEREMRRFAGACRFVFNRALALQNENHEAGNKYIPYGKMASWLVEWKNATETQWLKDAPSQPLQQSLKDLERAYKNFFRKRAAFPR FKKRGQNDAFRYPQGVKLDQENSRIFLPKLGWMRYRNSRQVTGVVKNVTASQSCGKWYISIQTENEVSTPVHPSALMVGLDAGVAKLATLSDGTVFGPVNSFQ KNQKTLARLQRQLSRKVKFSNNWQKQKRKIQRLHSCIANICRDYLHKVTTTVSKNHAMIVIEDLKVSNMSKSAAGTVSQPGRNVRAKSGLNRSILDQGWYEMRRQLEYKQLWRGGQVLAVPPAYTSQRCACCGHTAKENRLSQSKFRCQACGYTANADVNGARNILAAGHAVLACGEMVQSGRPLKQEPTEMIQATA(Sequence ID: 5)
[0065] IS1341 (IS1341 family) TnpB protein. Length: 369; NCBI accession number D38778 IS Finder Database Entry: https: / / isfinder.biotoul.fr / scripts / ficheIS.php?name=IS1341 MANKAYQFRLYPTKEQEQLLAKTFGCVRFVYNKMLEERIQMFEKFKDDQESLKQQTCPTPAKYKKEFPWLKEVDSLALANAQLNLQKAFQHFFSGRAGFPKFKNRKAKQSYTTNMVNGNIKLSDGYIKLPKLKWIKLKQHREIPAHHIIKSCTITKTKTGKYYISILTEYEHQPAPKEVQTVVGLDFSMSTLYVDSEGKRANYPRFYRKALETLAKEQRKWSRKKKGSNRWHKQRLKVAKLHEKIANQRKDFLHKESHKLAKRYDCVVIEDLNMKGMSQALHFGQGVHDNGWGMFTTFLQYKLVEQGKKLIKIDKWFPSSKTCSCCGRVKESLSLSERTFRCECGFESDRDVNAAINIKHEGMKRLAIV(Sequence ID:6)
[0066] ISC1316 (IS1341 family) TnpB protein. Length: 393; NCBI accession number: NC_002754 IS Finder Database Entry: https: / / isfinder.biotoul.fr / scripts / ficheIS.php?name=ISC1316 MPTLGFRFRAYTDEQTLRALKAQLKLTCEIYNTLRWADIYFYQRDGKGLTQTELRQLALDLRKQDDEYKQLYSQVVQQVADRYSEAKKRFFEGLARFPKE KKPHKYYSLVYTQSGWKILHVREIRKGKKNKKKLITLKLSNLGTFKVIVHRDFPLDKVKRVVVKLTRSERIYITFVVDHEFPKLPNTGKVVAIDVGVEKL LITSDGEYFPNLRPYEKALWKVKHIHRELSRKKFLSNNWFKAKVKLARAYEHLKNLRTDLYMKLGKWFAEHYDVVVMEGIHAKQLVGKSLRSLRRRLSDVGFGELRGVLKYQLEKYGKKLILVNPAYTSKTCARCGYVKNDLSLSDRVFVCPNCGWIADRDYNASLNILRGSGSERPLVWSSALYQYSGKVGL(Sequence ID:7)
[0067] IS891 (IS1341 family) TnpB protein. Length: 401; NCBI accession number M24855 IS Finder Database Entry: https: / / isfinder.biotoul.fr / scripts / ficheIS.php?name=IS891 MLVFETKLEGTNEQYQLLMRRLKLLVLSNACLRTWIGQPNIGRYDLSAYCAVLLPMKTFRSLPNSTLWLDKLLLKERGVQLLGFLTIASKTKPGRKVIHALK KNRRMGVLSIKLAAGSLVVTVAYVTFSDGFKAGTFKLWGTRDLHFYQLKQFKRVRVVRRADGYYAQFCIDQERVERREPTLKTIGLDVGLNHFLTDSEGNTV ENPRHLRKSEKSLKRLQRRLSKTKKGSNNRVKARNRLSRKHLKVSRQRKDFAVKLARCVVQSSDLVAYEDLQVRNMVRNRHLAKSISDAAWTQFRQWVEYFG KVFGVVTVAVPPHTSQNCSNCGEVVKKSLSTRTHACPHCGHIQDRDWNAARNILELGLRTVGHTGSQVSGDIDLCLGEVTPPNKSSRGKRKPKK (SEQ ID NO: 8)
[0068] ISEc42 (IS1341 family) TnpB protein. Length: 376; NCBI accession number NC_004431 IS Finder database entry: https: / / isfinder.biotoul.fr / scripts / ficheIS.php?name=ISEc42 MKRAYKYRFYPTTEQAELLAQTFGCVRFVYNSILRWRTDAYYERKEKIGYLQANARLTALKKEPEYIWLNDVSCVPLQQSLRHQQAAFANFFAGRAAYPAFKSKRHKQVAEFTASAFKHRDGELYIAKSKSPLDVRWSRELPSAPSTVTISRDSAGRYFVSCLCEFEPVSMPVTAKTVGIDVGLKDLFVTDTGFKTDNPRHTAKYAKRLTLLQRRLSRKQKGSRNRIKARLKVARLHAKIADCRMDNLHKLSRKLINENQVVCVESLKVKNMIRNPKLSKAIADAGWSELVRQLQYKGKWAGRSVVAIDQYLPSSKCCSCCGFTMQKMPLNVRKWHCPECGADHDRDINAARNIKAAGLAVLAHGEPVNPESQHAA(Sequence ID:9)
[0069] ISTel3 (IS1341 family) TnpB protein. Length: 393, NCBI accession number NC_004113. IS Finder Database Entry: https: / / isfinder.biotoul.fr / scripts / ficheIS.php?name=ISTel3 MRGVEKAFSYRFYPTTEQESLLRKTLGCVRLVYNRALAARTEAWYERKERLDYVQTSALLTQWKKQDDLQFLNEVSCVPLQQALRHLQSAFTNFFAGRAKYPNFKKKRNGGSAEFTKSAFRWKDGKVFLAKCNEPLNIPWSRRLPDGVEPSTVTIRLNPAGQWYISLRFDDPRELTLQPVDPSVGLDVGMSSLITLSTGEKIANPKHFNRYYKRLRKAQRSLSRKQKGSRNWDKARLKVAKIHQKISDSRKDHLHQLTTRLIRENQTIIIESLAVKNMVKNRQLARSISDAGWGELVRQLEYKAQWYGRTLVKIDRWFPSSKRCGQCGHIVEWLPLSVREWDCPKCGAHHDRDINAAGNILAVGHTVTVCGAGVRPDRHTSGGQLRRNRKSQK(Sequence ID: 10)
[0070] IS607 (IS607 family) TnpB protein. Length: 419; NCBI accession number AF189015 IS Finder Database Entry: https: / / isfinder.biotoul.fr / scripts / ficheIS.php?name=IS607 MSAISITHKIALKPNNKHITYFKKAFGCARFAYNWGLAKWKENYQLGIKTSHLQLKKEFNALKKSQFNFVYEVTKYATQQPFIHLNLAFNKFFRDLEKGLVSYPKFK KKREFQGSFYIGGDQIKIIQTANTDYLKIPNLPPIKLTEKLRFQGKIHNATITQKGDHFYVSISCDIDESEYKRTHKLQESHNKLGIDIGIKSFVSLSNGLNIYAPK PLDKLTRKLVRISRQLSKKIHPKTKGDKTRKSNNYLKHSKKLTHLHEKIANIRLDFLHKLTSSLIRHSNSFCLESLKVKNMFKNHRLAKSLSDISMSVFNTLLEYKAKYSNKEILRADTYYPSSKTCSNCQKVKQDLKLKDRIYQCLECGFELDRDINAAINLLKHLVGRVTAEFTPMDLTALLNDLSNNRLATSKVELGIQQKS (Sequence ID: 11)
[0071] ISTsi1 (IS607 family) TnpB protein. Length: 393; NCBI accession number NC_012883 IS Finder database entry: https: / / isfinder.biotoul.fr / scripts / ficheIS.php?name=ISTsi1 MPSETIKLASKFKLKETPEGLNELFSTYRDIVNFLITHAFENNITSFYRLKKEIYKSLRKEYPELPSHYIYTACQMAASIYKSYRKRKRRGKASGRPVFKKEAIMLDDHLFKLDLEKGIIKLSTPNGRITLKFYPAKHHEKFKNWKVGQAWLVRTPKGVFINVVFSKEVEVKEPEDFVGVDLNENNVTLSLSDGEFVQIITHEKEIRTGYFVKRRKIQKKVKVGKKRQELLEKYGERERNRLNDLYHKLANKIVELAEKYGGIALEDLTEIRNSIRYSAEMNGRLHRWSFRKLQSIIEYKAKLKGVEVVFVDPAYTSSLCPVCGEKLSPNGHRVLKCLNCGFEADRDVVGSWNVRLRALKMWGVSVPPESPPMKMGGGKASRGDVYELYTNYG (Sequence ID: 12)
[0072] IS1535 (IS607 family) TnpB protein Length: 550; from NCBI accession number Z95210 IS Finder Database Entry: https: / / isfinder.biotoul.fr / scripts / ficheIS.php?name=IS1535 (Sequence ID: 13)
[0073] ISBlo12 (IS607 family) TnpB protein Length: 440; from NCBI accession number NC_004307 IS Finder Database Entry: https: / / isfinder.biotoul.fr / scripts / ficheIS.php?name=ISBlo12 MSAYEAVRIRLDPTPRQTRLLESHAGGARFAYNLMLAHVRRQISLGEKPDWTLYAMRRWWNEWKDEIAPWWRENSKEAYGSAFEWLSQALRNWSDSRKGRRAGRRVGWPKYK SKRSSVPRFAYTTGSFGLIEDDPKALRLPRIGRVHCMENATERVHGRRIVRMTVSRHAGFWYAALTVERPTESVPAKNRKRKNHDRQVGVDLGVRTLATLSDGTTFPNPRNY VRTQRKLRHAQQSLSRRDRGMSHGCGSKRYNRALERVRRIHARIAAQRADNIGKLTTWLADNYSDISIEDLNVQGMSHNRRLAKHILDADFHEFRRQLEYKTARAGTRLHVIDRWYPSSKTCSNCGTVKAKLSLSERVYHCEECGLVIDRDVNAAINIQVAGSAPETLNARGGSVGQTRLECGTMRHPAKREPSGGDSRVRLGAGLGNEAMQMTSL (Sequence ID: 14)
[0074] ISC1926 (IS607 family) TnpB protein. Length: 412; NCBI accession number AY671948 IS Finder Database Entry: https: / / isfinder.biotoul.fr / scripts / ficheIS.php?name=ISC1926 MERTIKLRVRVDYITYSALKEVEGEYREVLEDAINYGLSNKTTSFTRIKAGVYKTEREKHKDLPSHYIYTACEDASERLDSFEKLKKRGRSYTEKPSVRKVTVHL DDHLWKFSLDKISISTMQGRVFISPTFPKIFWRYYNTEWRIASEARFKLLKGNVVEFFIVFKRDEPKPYEPKGFIPVDLNEDSVSVLVDGKPMLLETNTKRITLG YEYRRKAITTRRSAEDREVKRKLKRLRERDKKVVIRRKLAKLIVKEAFESMSAIVLEALPRRPPEHMIKDVKDSQLRLRIYRSAFSSMKNAIIEKAKEFRVPVVLVNPSYTSSTCPIHGAKIVYQPDGGDAPRVGVCEKGKEKWHRDVVALYNLRKRAGDVSPVPLGSKESHDPPTVKLGRWLRAKSLHSIMNEHKMIEMKV(Sequence ID: 15)
[0075] The TnpB protein-containing protein may additionally include one or more effector molecules, in particular, one or more effector molecules that are covalently linked to the TnpB protein to form a fusion protein. The fusion protein according to this disclosure is discussed further below.
[0076] This disclosure also relates to DNA and RNA encoding proteins, including the TnpB protein described herein, which are essentially composed of or consist of such proteins, and which can be produced therefrom by expression. DNA or RNA expression may occur in vitro, ex vivo, or in vivo.
[0077] Inactive TnpB protein The proteins used in the effector complex may also include, essentially be, or be derived from, a mutant TnpB protein in which its nuclease activity is partially or completely inactivated. Such proteins have one or more mutations in the RuvC-like domain of the protein that affect the nuclease activity of TnpB. In particular, point mutations in the RuvC-like domain that remove nuclease activity are already known in the art and have been used to generate the mutant Cas12(Cpf1). Mutations D917A and E1006A of FnCpf1 completely inactivate the cleavage activity of FnCpf1, while mutation D1225A has been reported to significantly reduce nuclease activity (Zetsche et al., 2015). Mutations of similar key residues in the RuvC-like domain of the TnpB protein may also be used to remove the nuclease function of TnpB and generate the inactivated mutant TnpB proteins described herein. As described above and shown in Figure 12, the RuvC-like domain of the TnpB protein typically contains a conserved D---E---D motif that can be mutated. The locations of these residues in each of SEQ ID NOs: 1-15 are shown in Figure 12. For example, in SEQ ID NO: 1 (TnpB protein from ISDra2), these are D191, E278, and D361. Equivalent residues in other TnpB proteins can be identified using sequence alignment tools, such as the Clustal Omega sequence alignment program (https: / / www.ebi.ac.uk / Tools / msa / clustalo / ) (Madeira et al., 2019).
[0078] Accordingly, in one example, an inactive mutant TnpB protein may include the TnpB proteins described herein, which have mutations in amino acid residues in the RuvC-like domain such that the nuclease activity is inactivated or partially inactivated. In particular, the mutations may be one, two, or three amino acid residues in the conserved D---E---D motif.
[0079] In certain cases, the mutant TnpB protein has a sequence having at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 1, wherein the sequence is mutated at at least one of positions D191, E278, and D361 of SEQ ID NO: 1 such that the RuvC-like domain is inactivated or partially inactivated.
[0080] MIRNKAFVVRLYPNAAQTELINRTLGSARFVYNHFLARRIAAYKESGKGLTYGQTSSELTLLKQAEETSWLSEVDKFALQNSLKNLETAYKNFFRTVKQSGKKVGFPRFR KKRTGESYRTQFTNNNIQIGEGRLKLPKLGWVKTKGQQDIQGKILNVTVRRIHEGHYEASVLCEVEIPYLPAAPKFAAGVDVGIKDFAIVTDGVRFKHEQNPKYYRSTLKR LRKAQQTLSRRKKGSARYGKAKTKLARIHKRIVNKRQDFLHKLTTSLVREYEIIGTEHLKPDNMRKNRRLALSISDAGWGEFIRQLEYKAAWYGRLVSKVSPYFPSSQLCHDCGFKNPEVKNLAVRTWTCPNCGETHDRDENAALNIRREALVAAGISDTLNAHGGYVRPASAGNGLRSENHATLVV (Sequence ID: 1, positions D191, E278, and D361 shown in bold) In other examples, the mutant TnpB protein has a sequence that has at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% sequence identity with one of sequence numbers 2–15, where the sequence is mutated at at least one of the amino acid residues enclosed in the box shown in Figure 12 such that the RuvC-like domain is inactivated or partially inactivated.
[0081] Effector complexes containing inactive TnpB proteins can be used simply to block specific target regions, including target sites in polynucleotides, for example, to disrupt transcription in those regions. They can also be used to detect the presence of polynucleotides containing target sequences in a sample, for example, in a manner in which the binding of the effector complex to the target site produces a measurable change in the physical or chemical properties of the detection system (e.g., in the context of biosensors).
[0082] Inactive TnpB proteins may also be used in effector complexes comprising one or more effector molecules. In these embodiments, the TnpB protein acts as a carrier for one or more effector molecules (which may also be called "one or more cargo molecules") to deliver them to a specific target region in a polynucleotide. In one example, one or more effector molecules (in particular when they are protein-based, e.g., enzymes or protein-labeled fluorescent proteins) may be "held" as a portion of a fusion protein with TnpB (which will be discussed further below). Alternatively, or additionally, one or more effector molecules may be "held" as a portion of RNA or bound to RNA, which will be discussed further below.
[0083] This disclosure also relates to DNA and RNA encoding proteins, including, essentially composed of or consisting of, the inactive TnpB protein described herein, the protein which can be produced therefrom by expression. DNA or RNA expression may occur in vitro, ex vivo, or in vivo.
[0084] TnpB fusion protein As described above, the effector complex may hold one or more effector molecules in the form of a fusion protein with the TnpB protein. In such an example of the present disclosure, the effector complex protein comprises the TnpB protein and one or more effector molecules fused to the N or C terminus of the TnpB protein. Depending on the desired function of the fusion protein in the effector complex, the fusion protein may comprise the TnpB protein or an inactive (mutant) TnpB protein identified above that does not contain an active nuclease domain.
[0085] One or more effector molecules may be nuclear localization signals (NLSs) that assist in the transport of proteins into the nucleus of a cell via nuclear transport. In particular, such NLSs may be used when the target polynucleotide is in the nucleus of a cell. Typically, an NLS is a short sequence of positively charged lysine or arginine present at or near the N or C terminus of a protein so that they are exposed on the protein surface when the protein complexes with RNA. Non-limiting examples of NLSs include the sequence PKKKRKV (SEQ ID NO: 18) from the SV40 large T antigen, and the bipartite NLS of nucleoplasmin containing basic amino acid KR and four K residues in two clusters separated by a spacer of about 10 amino acids (e.g., KRPAATKKAGQAKKK - SEQ ID NO: 19). Other NLSs are known in the art.
[0086] Depending on how the effector complex is delivered to the cell, the fusion protein may also contain a cell membrane permeable peptide (a short peptide that facilitates the uptake of the fusion protein into the cell).
[0087] Furthermore, or alternatively, the fusion protein may contain one or more effector molecules. These one or more effector molecules may be one or more effector molecules capable of modifying polynucleotides in the target region; one or more effector molecules that are trans-acting factors capable of increasing or decreasing transcription in the target region; and / or one or more effector molecules capable of labeling the target region.
[0088] Methods for delivering one or more effector molecules to a target region using Cas9 and Cas12 fusion proteins are already known in the art (e.g., described in Knott et al., 2018 and Anzalone et al., 2020). Similar components can be fused to TnpB proteins or inactive TnpB proteins. Small-sized TnpB proteins, in particular, serve as good scaffolds for the generation of fusion proteins.
[0089] In particular, one or more effector molecules may be selected from endonucleases, ribonucleases, nickases, base editors, epigenetic modifiers, transposases, recombinases, and reverse transcriptases. Specifically, if the base editor is a deaminase, it may be cytidine deaminase and / or adenosine deaminase. A fusion protein containing cytidine deaminase may also contain a uracil glycosylase inhibitor.
[0090] One or more effector molecules may be used in the fusion protein for labeling the target region. The labeling may be a reporter enzyme such as GFP or a fluorescent protein, which can be used to detect the effector complex when the guide RNA hybridizes to the target sequence.
[0091] One or more effector molecules may be used in the fusion protein to increase or decrease transcription or translation in a target region. These may be one or more transcription activators or one or more transcription repressors.
[0092] This disclosure also relates to DNA and RNA encoding a fusion protein comprising a TnpB protein (or inactive TnpB protein) and one or more effector molecules described herein, the fusion protein of which may be produced by expression. DNA or RNA expression may occur in vitro, ex vivo, or in vivo.
[0093] RNA used in effector complex This disclosure also relates to RNA capable of binding to the TnpB protein to form an effector complex, and capable of guiding or inducing the effector complex to a target region in a polynucleotide.
[0094] In particular, this disclosure provides RNA including: (i) A protein-binding segment that enables RNA to bind to the TnpB protein and form an effector complex, and (ii) A polynucleotide targeting segment containing a guide sequence capable of hybridizing to a target sequence in the target region of a polynucleotide.
[0095] The protein-binding segment of the RNA interacts with the TnpB protein, binding the RNA to the TnpB protein and forming an effector complex. The protein-binding segment may contain sequences capable of forming RNA secondary structures. The protein-binding segment may contain at least one inverted repeat sequence, i.e., a sequence section whose reverse complementary sequence follows downstream, thereby allowing the two sections to hybridize to form a double-stranded RNA (dsRNA) duplex, such as a hairpin, an incomplete hairpin, or other secondary RNA structure. In particular, one or more inverted repeat sequences may be sequences that are at least partially palindromic structures, thereby allowing the sequences to form at least one hairpin or at least one incomplete hairpin (also referred to as a stem-loop or hairpin-loop).
[0096] The protein-binding segment may contain a sequence from the right end (RE) of the insertion sequence in the IS200 / IS605 or IS607 family (where thymine residues in the RE DNA sequence are replaced by uracil residues). The RE sequence may be an incomplete palindromic sequence from a mobile gene in the IS200 / IS605 family. The RE sequence may incorporate a portion of the terminal sequence of the tnpB gene. The RE sequence may be from the same mobile gene as tnpB, from which the TnpB protein in the effector complex originates. The RE sequence of a particular insertion sequence may be known in the art (for example, from the same insertion sequence as the TnpB protein with sequence numbers 1 to 15 referenced above, which may be available in the ISfinder database). Alternatively, the RE sequence may be determined based on sequencing of the right end of the insertion sequence that moves with the tnpB gene during transposition. The RE sequence section that can be used in the protein-binding segment can be determined in an assay in which the tnpB gene is co-expressed in a suitable host cell (such as Escherichia coli) having the complete insertion sequence (optionally, having an inactivated tnpA gene present in the insertion sequence), and subsequently, the TnpB-binding RNA is characterized by small RNA sequencing, as described in Example 1 herein, for example.
[0097] In one example, the protein-binding segment contains or consists of SEQ ID NO: 16, GAAUCACGCGACUUUAGUCGUGUGAGGUUCAA (which can form an incomplete hairpin as shown in Figure 1D). This sequence is from the RE of the insert sequence ISDra2. Therefore, preferably, if the protein described above contains a TnpB protein with the amino acid sequence of SEQ ID NO: 1 (which is from the tnpB gene of ISDra2), the protein-binding segment of the RNA contains or consists of SEQ ID NO: 16.
[0098] The polynucleotide targeting segment of RNA contains a guide sequence that is complementary to, or capable of hybridizing to, the target sequence in the target region of the polynucleotide. This segment of RNA acts to induce or target the effector complex to the target region in the polynucleotide.
[0099] The target sequence into which the RNA hybridizes may be single-stranded DNA or part of a double-stranded DNA polynucleotide. In examples where the effector complex contains a mutant / inactive TnpB and is used to block a target region or to deliver one or more effector proteins to a target region, the target sequence into which the DNA hybridizes may be RNA. (As described herein, if the polynucleotide containing the target sequence is double-stranded DNA, the site where site-specific cleavage of the polynucleotide occurs is determined by both complementary base pairs between the guide sequence and the target sequence, and a short TnpB association sequence motif (TAM) that interacts with the TnpB protein.) The RNA guide sequence may be 10–30 nucleotides or 15–25 nucleotides long and have sufficient complementarity to the target sequence to enable hybridization between the guide sequence and the target sequence under the specific conditions in which the effector complex is used. In most situations, a complementarity of 80% or higher is preferred.
[0100] The two segments of the RNA are covalently linked as a single RNA molecule, and optionally, there may be intervening linker ribonucleotides that separate the two segments. The RNA may consist of a 5' protein-binding segment - (optional linker) - polynucleotide targeting segment - 3' or a 5' polynucleotide targeting segment - (optional linker sequence) - protein-binding segment - 3'. Preferably, the composition is a 5' protein-binding segment - (optional linker) - polynucleotide targeting segment - 3'.
[0101] Overall, RNA can be 50–300 nucleotides long, 100–200 nucleotides long, or 140–150 nucleotides long.
[0102] It should be noted that RNA is a designed RNA that does not occur naturally; that is, RNA is artificially produced, and neither the polynucleotide targeting segment nor the protein-binding segment occurs naturally.
[0103] In particular, in a preferred embodiment, the guide RNA is complementary to non-bacterial, non-archaeal gene sequences.
[0104] The RNA provided in this disclosure may include, for example, chemical modifications to reduce RNA degradation in target cells. Techniques for testing modifications in crRNA and tracrRNA used in CRISPR Cas9 and Cas12 systems have already been described in the art and may be applied (e.g., Mir et al., 2018). The RNA molecule may further include a segment that allows the RNA to bind to one or more effector molecules delivered to a target region containing a polynucleotide target sequence. Aptamers such as MS2 hairpins or PP7 hairpins may be designed within the RNA to which effector molecules (e.g., MS2 RNA-coated protein MCP fused to a fluorescent protein) may be anchored or bound in the manner described in the art for dCas9, for example (Sajwan S, et al., 2019; Ma H, et al., 2018; Ma et al., 2016).
[0105] This disclosure also relates to DNA encoding RNA as described herein, the RNA which can be produced therefrom by expression. DNA expression may occur in vitro, ex vivo, or in vivo.
[0106] Effector complex Effector complexes comprising the proteins and RNAs identified above are also provided by this disclosure. These are guided by RNA to a target sequence in the target region of a polynucleotide, and the RNA includes a polynucleotide-binding segment containing a guide sequence that hybridizes to the target sequence of the polynucleotide.
[0107] The polynucleotide to which the effector complex is induced may be double-stranded DNA or single-stranded DNA. Preferably, the polynucleotide is double-stranded DNA. In examples where the effector complex contains a mutant / inactive TnpB and is used to block a target region or to deliver one or more effector proteins to a target region, the target sequence to which the effector complex is induced may be RNA.
[0108] If the effector complex contains a TnpB with an active nuclease site, the effector complex can cleave DNA in the target region. The cleavage may occur within 30 bp from the edge of the target site. The cleavage site may be 5' of the target sequence on the strand containing the target sequence.
[0109] In one example, the effector complex can cleave a double-stranded polynucleotide to produce a stepped double-stranded cleavage region. The 5' overhang may be, for example, 4 or 5 nucleotides long. Alternatively, the effector complex can cleave a double-stranded polynucleotide to produce a blunt end.
[0110] The effector complexes of this disclosure may be engineered, non-naturally occurring complexes. In particular, both the RNA and protein of the complexes do not occur naturally.
[0111] The effector complex may be in an isolated or purified form.
[0112] In one example of this disclosure, the effector complex is bound to a solid support. In particular, the effector complex may be bound to a solid support in a biosensor that can be used to detect the presence of a target sequence (e.g., Hajian et al., (2019) describes a Cas9-based effector complex immobilized on a graphene field effector transistor). Preferred methods for attaching proteins to solid surfaces are known in the art, which can be used to attach the effector complex to the solid surface. In one example, the effector complex may comprise the fusion protein described above, comprising a TnpB protein (or an inactivated TnpB protein) and a peptide tag that can be used to capture the effector complex on the surface of the solid support.
[0113] The effector complexes of this disclosure may be produced in vitro, ex vivo, or in vivo. In particular, the method may include assembling the effector complex from the RNA and proteins described herein in vitro in cells or in a cell-free system.
[0114] If the effector complex is produced in cells, the method may include providing the following in cells: (i) DNA encoding RNA and proteins as described herein; (ii) DNA encoding proteins and RNAs as described herein; (iii) DNA encoding proteins described herein, and DNA encoding RNA described herein; (iv) RNA (mRNA) encoding a protein as described herein, and RNA as described herein; or (v) RNA (mRNA) encoding a protein as described herein, and DNA encoding RNA as described herein.
[0115] When the effector complex is produced in vitro in a cell-free system, the method may include in vitro expression of RNA-coding DNA, in vitro expression of protein-coding DNA, or in vitro expression of both RNA-coding DNA and protein-coding DNA.
[0116] Protein-coding DNA and / or RNA-coding DNA may contain one or more regulators that regulate DNA expression in cells or cell-free systems. In particular, protein-coding DNA may contain at least one primary regulator operably ligated to the protein-coding DNA sequence, and / or RNA-coding DNA may contain at least one secondary regulator operably ligated to the RNA-coding DNA sequence. "Operatally ligated" means that the regulator is positioned in the DNA sequence in such a way that it can influence the expression of RNA and protein-coding DNA sequences. Regulators may be promoters, enhancers, internal ribosome entry sites, and other expression regulators. These may be selected depending on the cell type used to express RNA and proteins, or other components selected for use in in vitro cell-free systems.
[0117] The DNA sequences disclosed herein may be incorporated into vectors. In particular, vectors may be used for the expression, maintenance, and / or propagation of DNA sequences. Preferred vectors include plasmids and viral vectors. Viral vectors may be selected from retroviral vectors, lentiviral vectors, adenovirus vectors, adeno-associated virus (AAV) vectors, or herpes simplex virus vectors. In particular, viral vectors already known in the art for use in combination with CRISPR-Cas9 and CRISPR-Cas12 systems may be used (e.g., described in Xu et al., 2019). In a preferred example, the viral vector is an AAV viral vector. In particular, due to the relatively small size of the TnpB protein, AAV viral vectors may be used especially when TnpB (or inactivated TnpB) for the effector complex is a portion of a fusion protein that holds one or more effector molecules.
[0118] This disclosure also provides host cells transfected with RNA-coding DNA and / or protein-coding DNA as described herein. The host cells may be used for the in vitro expression of RNA-coding DNA and / or protein-coding DNA as described herein, and in particular for the generation of effector complexes. The host cells contain RNA-coding DNA and / or protein-coding DNA as described herein. The DNA may be incorporated into the genome of the host cell so as to be replicated along with the host genome. Alternatively, the DNA may remain on the vector used to transfect the cells.
[0119] DNA can be defined as something foreign to the host cell; that is, host cells containing DNA do not arise naturally.
[0120] In some cases, the host cells are isolated cells.
[0121] In some cases, the host cells are not totipotent human embryonic stem cells.
[0122] In some cases, the host cell is not a human oocyte.
[0123] In some cases, the host cell does not contain a target sequence complementary to the RNA guide sequence.
[0124] The host cells may be cells from a cell line.
[0125] In one aspect of this disclosure, a host cell can be used to produce the effector complex described herein, which can then be used in the manner described below.
[0126] In alternative embodiments of this disclosure, the generation of the effector complex may occur as part of the methods and uses of the effector complex discussed herein.
[0127] Both of these embodiments may involve the following systems, which are also provided by this disclosure.
[0128] A system for modifying a target region in a polynucleotide, where the target region includes a target sequence, and the system is: a) A protein containing or consisting of the TnpB protein, or DNA or RNA encoding said protein, and b) RNA, or RNA-coding DNA It includes RNA: (i) a polynucleotide targeting segment containing a sequence complementary to the target sequence; and (ii) Protein-binding segment that binds to TnpB protein Includes.
[0129] The system may include (a) a protein and (b) RNA; (a) DNA encoding a protein and (b) DNA encoding RNA; (a) DNA encoding a protein and (b) RNA; (a) DNA encoding a protein and (b) RNA; (a) RNA encoding a protein (mRNA) and (b) RNA; or (a) RNA encoding a protein (mRNA) and (b) DNA encoding RNA. In certain examples, both (a) and (b) are RNA (as shown for Cas9, e.g. (Gillmore et al., 2021)).
[0130] The RNA and TnpB-containing proteins are as described herein. In particular, the proteins may be fusion proteins as described herein. The TnpB may be inactivated TnpB as described herein.
[0131] In the system, (a) and / or (b) may be contained in at least one vector. In one example, (a) and (b) are contained in the same vector. In an alternative example, (a) and (b) are contained in separate vectors.
[0132] The vector can be a non-viral vector of a viral vector. In particular, a non-viral vector can be at least one plasmid and / or at least one non-viral particle such as a liposome or exosome.
[0133] Alternatively, at least one viral vector may be selected from retroviral vectors, lentiviral vectors, adenovirus vectors, adeno-associated virus (AAV) vectors, or herpes simplex virus vectors. As described above, an AAV vector may be preferred in particular if the protein is a fusion protein comprising TnpB and one or more effector molecules.
[0134] The system described in this disclosure is designed to be a system that does not occur naturally.
[0135] The system may be in the form of a kit in which (a) and (b) are packaged separately, and optionally the kit may be packaged with instructions for use.
[0136] The systems and effector complexes described above may be contained in the vectors described above for in vitro, ex vivo, or in vivo delivery to cells.
[0137] Furthermore, the system or effector complex may be delivered either alone or as part of a vector, by microinjection or via electroporation. In particular, the vector may be a liposome.
[0138] The system or effector complex can be delivered chemically via lipofection (mediated by lipids), transfection (mediated by cationic polymers), or calcium phosphate transfection.
[0139] Viral vectors, including lentiviral vectors, retroviral vectors, and AAV vectors, can also be used for delivery.
[0140] In particular, delivery systems based on those already described for CRISPR Cas9 and Cas12 systems can be used (see, for example, https: / / blog.addgene.org / crispr-101-mammalian-expression-systems-and-delivery-methods).
[0141] As described above, the fusion protein used in the effector complex may include a cell membrane-permeable peptide to facilitate the uptake of the effector complex or fusion protein by cells.
[0142] Method and Use The effector complexes and / or systems described herein may be used in methods for cleaving, modifying, labeling, or controlling the expression of a target region in a polynucleotide, wherein the target region includes a target sequence.
[0143] In particular, the method may be a method for delivering an effector complex to a target region in a polynucleotide, where the target region includes a target sequence, and the effector complex includes: (a) Proteins containing or consisting of TnpB protein; and (b) RNA containing the following: (i) a polynucleotide targeting segment containing a guide sequence capable of hybridizing to a target sequence; and (ii) A protein-binding segment that enables RNA to bind to the TnpB protein. The method comprises delivering the effector complex to a target region by contacting the effector complex with a polynucleotide, thereby enabling the guide sequence to hybridize with the target sequence. The effector complex may comprise one or more effector molecules as described herein, which are delivered to the target region.
[0144] In a further embodiment, the method may be a method for cleaving a polynucleotide using an effector complex, where the polynucleotide comprises a target sequence and the effector complex comprises: (a) Proteins containing or consisting of TnpB protein; and (b) RNA containing the following: (i) a polynucleotide targeting segment containing a guide sequence capable of hybridizing to a target sequence; and (ii) A protein-binding segment that enables RNA to bind to the TnpB protein to form an effector complex, wherein the method comprises contacting a polynucleotide with the effector complex and enabling the TnpB protein to cleave the polynucleotide.
[0145] If the polynucleotide is double-stranded DNA, the cleavage may produce a stepped double-strand break with a 5' overhang. Alternatively, the cleavage may produce a blunt-end double-strand break.
[0146] The contact step of the method may occur in cells under conditions that allow non-homologous end joining (NHEJ) or homologous recombination repair (HDR) of the cleaved polynucleotides, thereby editing the polynucleotide sequence. Furthermore, the method may further comprise a step of contacting the polynucleotide with a donor polypeptide for HDR. Preferred methods for achieving NHEJ and HDR known in the art for Cas9 and Cas12 systems are also preferred in this case (see, for example, Maresca et al., 2013).
[0147] In these embodiments of the method, the polynucleotide may be double-stranded DNA and may contain a TnpB association sequence motif 5' of a target sequence (described above) with which TnpB interacts.
[0148] Alternatively, polynucleotides can be single-stranded DNA.
[0149] In one example of the method of this disclosure, polynucleotides may be located inside a cell. The cell may be a prokaryotic or eukaryotic cell. If the cell is a eukaryotic cell, it may be a non-human animal cell, a human cell, or a plant cell. In particular, the cell may be a stem cell, such as an induced pluripotent stem cell.
[0150] Similar to the Cas9 and Cas12 systems, the methods of the present disclosure are particularly useful in plant cells. In particular, the present disclosure includes a method for producing a plant comprising cells having a modified polynucleotide, the method comprising the step of contacting a plant cell with the system or effector complex described herein to modify a target region of the polynucleotide, thereby regenerating a plant from the plant cell, wherein the modified target region is in a gene of interest in the cell, and the modification relates to a property of interest.
[0151] Effector complexes, systems, and the DNA encoding components of the complexes and systems may be used as pharmaceuticals in individuals. Alternatively, they may be used for diagnostic methods in individuals.
[0152] Alternatively, effector complexes, systems, and DNA encoding components of the complexes and systems may be used in vitro or ex vivo to determine the presence of polynucleotides containing target sequences in a sample, or to modify target regions of polynucleotides.
[0153] Here, this disclosure will be described in more detail by referring to the following experimental procedures, simply as an example.
[0154] example material and method Design of a TnpB expression vector The pTWIST-ISDra2 plasmid containing the IS200 / IS605 ISDra2 system of Deinococcus radiodurans R1 (GenBank AE000513.1), cloned as a synthetic DNA fragment under the T7 promoter, was obtained from Twist Biosciences. To obtain a pGD3 plasmid containing an ISDra2 variant with a deletion within the tnpA gene, the pTWIST-ISDra2 plasmid was pre-cut with NdeI (Thermo Fisher Scientific), the 5'-overhang was embedded using T4 DNA polymerase (Thermo Fisher Scientific), and autocirculated using T4 DNA ligase (Thermo Fisher Scientific). For TnpB purification, two pBAD-derived expression vectors were constructed using the NEBuilder HiFi DNA Assembly Kit (New England Biolabs). pTK120-ISDra2-TnpB contains a tnpB coding sequence fused to an N-terminal 10×His TwinStrep-MBP protein purification tag, while pTK151 contains a tnpB coding sequence fused to an N-terminal 6×His-MBP and a C-terminal StrepTag II coding sequence. To obtain the reRNA expression vector (pGB71) used for TnpB complex purification, an reRNA coding sequence holding the T7 promoter at the 5' end, and HDV (hepatitis D virus) ribozyme and T7 terminator (assembled by PCR from synthetic oligonucleotides) at the 3' end were cloned into the pACYC184 vector on HindIII and BclI restriction sites (Thermo Fisher Scientific). The pGB74-78 plasmids used for TnpB complex expression and plasmid interference assays in 7N plasmid library cleavage contain reRNA and tnpB coding sequences under the T7 and T7lac promoters, respectively.The pGB74-78 plasmid was obtained by cloning reRNA coding fragments on the Bsu15I and EcoRI (Thermo Fisher Scientific) sites, and tnpB on the NdeI and XhoI (Thermo Fisher Scientific) sites, into the pET-Duet1 vector (Novagen). For genome editing experiments in human HEK293T cells, plasmid vectors pRZ122-127, derived from the pX458 plasmid (donated from Feng Zhang, Addgene plasmid #48138), encoding reRNA (targeting a 20bp site in human genomic DNA) and tnpB (fused with SV40 NLS-T2A-GFP at the 3' end) under the U6 and CAG promoters, respectively, were constructed using the NEBuilder HiFi DNA Assembly Kit (New England Biolabs). The Phusion Site-Directed Mutagenesis Kit (Thermo Fisher Scientific) was used to obtain plasmid variants with mutant RuvC active sites.
[0155] Expression and purification of the TnpB RNP complex For initial TnpB protein expression and pre-purification, E. coli BL21-AI cells were transformed with pTK120-ISDra2-TnpB alone, or co-transformed with pGD3 (encoding the ISDra2 transposon with a deletion in the tnpA gene), and grown at 37°C in LB broth supplemented with ampicillin (100 μg / ml) or ampicillin (100 μg / ml) and chloramphenicol (50 μg / ml), respectively. OD 0.6-0.8 600After culturing, protein expression was induced with 0.2% arabinose, and the cells were grown for a further 16 hours at 16°C. The following day, the cells were pelleted by centrifugation and resuspended in a buffer containing 20 mM Tris-HCl, pH 8.0, 25°C, 250 mM NaCl, 5 mM 2-mercaptoethanol, 25 mM imidazole, 2 mM PMSF, and 5% (v / v) glycerol, and then disrupted by sonication. After removing cell debris by centrifugation, the supernatant was treated with Ni 2+ The proteins were loaded onto a HiTrap chelate HP column (GE Healthcare) and eluted in a linear gradient of increasing imidazole concentrations from 25 mM to 500 mM in 20 mM Tris-HCl, pH 8.0, 25°C, 500 mM NaCl, 5 mM 2-mercaptoethanol, and 5% (v / v) glycerol buffer. The fractions containing TnpB were pooled and dialyzed against 20 mM Tris-HCl, pH 8.0, 25°C, 250 mM NaCl, 2 mM DTT, and 50% (v / v) glycerol, and stored at -20°C. The obtained pre-purified TnpB samples were used for nucleic acid extraction and analysis.
[0156] To increase the expression and yield of the TnpB RNP complex, E. coli BL21-AI cells are used to express reRNA (pGB71) and TnpB (pTK151) or TnpB D191A (pTK152) transformed cells were grown in LB broth at 37°C supplemented with ampicillin (100 μg / ml) and chloramphenicol (50 μg / ml). OD values were 0.6–0.8. 600 After culturing, protein expression was induced with 0.2% arabinose, and the cells were grown for a further 16 hours at 16°C. The following day, the cells were pelleted by centrifugation and resuspended in a buffer containing 20 mM Tris-HCl, pH 8.0, 25°C, 500 mM NaCl, 5 mM 2-mercaptoethanol, 25 mM imidazole, 2 mM PMSF, and 5% (v / v) glycerol, and disrupted by sonication. After removing cell debris by centrifugation, the supernatant was treated with Ni 2+The reaction mixture was loaded onto a HiTrap chelate HP column (GE Healthcare) and the bound protein was eluted in a linear gradient of increasing imidazole concentrations from 25 to 500 mM in 20 mM Tris-HCl, pH 8.0, 25°C, 500 mM NaCl, 5 mM 2-mercaptoethanol, and 5% (v / v) glycerol buffer. The fraction containing the TnpB RNP complex was pooled, and the 6×His-MBP tag was cleaved using TEV protease by incubation overnight at 8°C. The reaction mixture was then loaded onto a StrepTrap column (GE Healthcare) and washed with 20 mM Tris-HCl, pH 8.0, 25°C, 150 mM NaCl, 5 mM 2-mercaptoethanol, and 5% (v / v) glycerol buffer, and the bound TnpB complex was eluted in 2.5 mM d-desthiobiotin solution. Fractions containing TnpB were pooled and loaded onto a HiTrap heparin HP column (GE Healthcare), and eluted using a linear gradient of NaCl concentrations increasing from 0.15 M to 1.0 M. The obtained TnpB complex fractions were pooled, concentrated to a maximum of 0.5 ml using an Amicon Ultra-15 centrifugal filter unit (Merck Millipore), and loaded onto a Superdex 200 10 / 300 GL (GE Healthcare) gel filtration column equilibrated with 20 mM Tris-HCl, pH 8.0, 25°C, 250 mM NaCl, and 5 mM 2-mercaptoethanol buffer. Peak fractions containing the TnpB RNP complex were pooled, dialyzed against 20 mM Tris-HCl, pH 8.0, 25°C, 250 mM NaCl, 2 mM DTT, and a buffer containing 50% (v / v) glycerol, and stored at -20°C. The concentration of the TnpB RNP complex was determined by quantifying the intensity of the protein band on an SDS-PAGE gel and comparing it to a protein standard of known concentration.
[0157] Molecular weight measurement by mass photometry The measurement coverslip (No. 1.5 H, 24 × 50 mm, Marienfeld) was purified for 5 minutes by continuous sonication in MilliQ water, isopropanol, and MilliQ water, and then dried using a stream of clean nitrogen gas. The clean coverslip was placed on a OneMP mass photometer (Refeyn Ltd.), and a CultureWell® reusable gasket (Grace Bio-Labs) was placed on top. The gasket well was filled with 10 μl of 20 mM Tris-HCl, pH 8.0, 25°C, and 250 mM NaCl buffer, and 10 μl of diluted TnpB RNP complex sample (approximately 60 nM) was added. The adsorption of biomolecules was monitored for 120 seconds using AcquireMP software (Refeyn Ltd). To convert the measured ratiometric contrast to molecular weight, the Un1Cas12f1 protein (Karvelis et al., 2020) and its oligomers (monomers or tetramers) in the range of 60–250 kDa were used for calibration. Samples were measured in triplicates. Mass photometry videos were analyzed using DiscoverMP (Refeyn Ltd).
[0158] TnpB-bound nucleic acid extraction and analysis To extract TnpB-bound nucleic acids, first, 100 μl of pre-purified TnpB sample was incubated with 5 μl (20 mg / ml) of proteinase K (Thermo Fisher Scientific) at 37°C for 45 minutes in 1 ml of 10 mM Tris-HCl pH 7.5, 37°C, 5 mM MgCl2, 100 mM NaCl, 1 mM DTT, and 1 mM EDTA reaction buffer. Next, the nucleic acids were extracted with a phenol:chloroform:isoamyl alcohol (25:24:1) solution, and the aqueous phase was further treated with chloroform to remove any residual phenol. The nucleic acid-containing solution was divided into fresh tubes (198 μl each), and then 2 μl of RNase I (10 U / μl) (Thermo Fisher Scientific) or DNase I (10 U / μl) (Thermo Fisher Scientific) was added, and the reaction was incubated at 37°C for 45 minutes. The reaction products were mixed with 2× RNA loading dye (Thermo Fisher Scientific), separated on a TBE-Urea (8M) 15% denatured polyacrylamide gel using 0.5× TBE electrophoresis buffer (Thermo Fisher Scientific), and visualized with SYBR® Gold (Thermo Fisher Scientific).
[0159] RNA isolation from TnpB RNP complex For TnpB-binding RNA extraction, 100 μl of pre-purified TnpB complexes were incubated with 5 μl (20 mg / ml) of proteinase K (Thermo Fisher Scientific) at 37°C for 45 minutes in 1 ml of reaction buffer containing 10 mM Tris-HCl, pH 7.5, 37°C, 5 mM MgCl2, 100 mM NaCl, 1 mM DTT, and 1 mM EDTA. The DNA was digested by adding 10 μl of DNase I (10 U / μl) (Thermo Fisher Scientific), followed by a further 45 minutes of incubation at 37°C, and then purified using the GeneJET RNA Cleanup and Concentration Micro Kit (Thermo Fisher Scientific). Next, 3 μg of purified RNA was phosphorylated in a 20 μl reaction volume at 37°C for 30 minutes using 1 μl (10 U / μl) of PNK (Thermo Fisher Scientific) in 1× reaction buffer A (Thermo Fisher Scientific) supplemented with 1 mM ATP, and then purified using the GeneJET RNA Cleanup and Concentration Micro Kit (Thermo Fisher Scientific).
[0160] RNA sequencing and analysis RNA libraries were prepared using the Collibri® Stranded RNA Library Prep Kit for the Illumina® System (Thermo Fisher Scientific) according to the manufacturer's instructions for small RNAs (protocol MAN0025359), pooled in equimolar ratios, and paired-end sequencing (2 × 75 bp) was performed using MiSeq Reagent Kit v2, 300 cycles (Illumina) on the MiSeq System (Illumina). Paired-end reads shorter than 20 bp were filtered with Cutadapt (Martin, 2011). Residual reads were mapped to a transposon-coding plasmid (pTWIST-ISDra2) using BWA (Li and Durbin, 2009) and converted to the .bam file format using SAMtools (Li et al., 2009). The resulting coverage data was visualized using IGV (Robinson, et al., 2011).
[0161] Detection of TnpB dsDNA cleavage and TAM recognition PAM determination assays previously developed for Cas9 and Cas12 effectors (Karvelis et al., 2015, 2019, 2020) were employed to establish TnpB dsDNA cleavage requirements and TAM sequences. Briefly, tnpB genes and reRNA constructs targeting 16bp or 20bp sequences in a plasmid library adjacent to the 7N randomization region were cloned into pET-duet1 (MilliporeSigma) vectors (pGB77-78). Next, E. coli Arctic Express (DE3) cells were transformed with plasmids encoding TnpB RNPs, and the cells were grown in LB broth supplemented with ampicillin (100 μg / ml) and gentamicin (10 μg / ml). 0.5 OD 600After reaching the target cell count, TnpB expression was induced using 0.5 mM IPTG, and the culture was incubated overnight at 16°C. Cells from 10 ml of the overnight culture were collected by centrifugation, resuspended in 1 ml of lysis buffer (20 mM phosphate, pH 7.0, 0.5 M NaCl, 5% (v / v) glycerol, 2 mM PMSF), and lysed by sonication. Cell debris was removed by centrifugation, and 10 μl of the supernatant containing TnpB RNPs was used directly for plasmid library digestion. Briefly, the lysate was mixed with 1 μg of a 7N randomized plasmid library (pTZ57) in 100 μl of reaction buffer (10 mM Tris-HCl, pH 7.5, 37°C, 100 mM NaCl, 1 mM DTT, and 10 mM MgCl2) and incubated at 37°C for 1 hour. The cleaved DNA ends were repaired by adding 1 μl of T4 DNA polymerase (Thermo Fisher Scientific) and 1 μl of 10 mM dNTP mix (Thermo Fisher Scientific), incubated at 11°C for 20 minutes, and then heated at a maximum of 75°C for 10 minutes each. Next, a 3'-dA overhang was added by incubating the reaction mixture with 1 μl of DreamTaq polymerase (Thermo Fisher Scientific) and 1 μl of 10 mM dATP (Thermo Fisher Scientific) at 72°C for 30 minutes. RNA was removed by adding 1 μl of RNase A (Thermo Fisher Scientific) and incubating the reaction mixture at 37°C for 15 minutes, and then DNA was purified using a GeneJet PCR purification kit (Thermo Fisher Scientific). Next, 100 ng of purified cleavage product was mixed with 100 ng of dsDNA adapter containing 3'-dT overhang (100 ng) and incubated with 1 μl of T4 DNA ligase (Thermo Fisher Scientific) in a 20 μl reaction volume at 22°C for 1 hour.Next, the cleavage products, including the adapters, were PCR amplified and gel-purified using the GeneJet gel purification kit (Thermo Fisher Scientific). The DNA library was prepared using the Collibri® PS DNA Library Prep Kit for the Illumina® system (Thermo Fisher Scientific) according to the manufacturer's instructions, pooled in equimolar ratios, and paired-end sequencing (2 × 150 bp) was performed using the MiSeq Reagent Kit v2, 300 cycles (Illumina) on the MiSeq System (Illumina).
[0162] Double-strand DNA breaks by the TnpB RNP complex were evaluated by investigating adapter ligation at target sequences in a 7N plasmid library. This was achieved by extracting and counting all reads containing adapters ligated at target positions 0–30 bp after the 7N region, by identifying 10 bp that perfectly match sequences derived from the adapter and plasmid backbone. Reads showing an increased frequency of adapter ligation in the target region (20–21 bp from the 7N randomized sequence) were used for 7N sequence (TAM) extraction and visualization using WebLogo (Crooks, 2004). The Python® scripts used for break site identification and TAM characterization are available in the GitHub repository (https: / / github.com / tkarvelis / Nuclease_manuscript).
[0163] DNA substrates for in vitro TnpB cleavage reaction The plasmid DNA substrates (pGB72-73) used in the in vitro cleavage assay were obtained by cloning a synthetic oligoduplex (Invitrogen) into a pSG4K5 plasmid (donated by Xiao Wang, Addgene plasmid #74492) that had been pre-cleaved with EcoRI and NheI restriction endonucleases (Thermo Fisher Scientific).
[0164] The synthetic linear DNA substrate contains 1 μM oligonucleotide (Thermo Fisher Scientific) and 1 μl (10 U / μl) PNK (Thermo Fisher Scientific) and 32 The oligoduplex (100 nM) was labeled at the 5' end by incubation with P-γ-ATP (PerkinElmer) in 7.5 μl of 1× reaction buffer A (Thermo Fisher Scientific) at 37°C for 30 minutes. 32 The compounds were obtained by combining P-labeled and unlabeled complementary oligonucleotides (1:1.5 molar ratio), which were then heated to 95°C and slowly cooled to room temperature.
[0165] DNA cleavage assay Plasmid DNA cleavage was initiated by mixing a 100 nM TnpB RNP complex with 3 nM plasmid DNA (pGB72-73) in a reaction buffer containing 10 mM MgCl2, 1 mM DTT, 1 mM EDTA, and 100 mM NaCl at 37°C, followed by incubation at 37°C for 60 minutes (unless otherwise indicated). The reaction was stopped by mixing with a 3 × loading dye solution (0.01% bromophenol blue and 75 mM EDTA in 50% (v / v) glycerol) and analyzed by agarose gel electrophoresis and ethidium bromide. The unwound plasmid DNA substrate was obtained by cleavage using NdeI endonuclease (Thermo Fisher Scientific).
[0166] The cleavage reaction using synthetic oligoduplexes was initiated in 100 μl of reaction buffer (100 mM NaCl, 10 mM MgCl2, 1 mM DTT, 1 mM EDTA, 100 μM Tris-HCl, pH 7.5) at 37° C. by combining 100 nM TnpB RNP complex with 1 nM radiolabeled substrate. 10 μl aliquots were removed from the reaction mixture at time intervals (0 min, 1 min, 5 min, 15 min, and 60 min), quenched with 1.8× volume of loading dye (95% (v / v) formaldehyde, 0.01% bromophenol blue and 25 mM EDTA), and subjected to denaturing gel electrophoresis (20% polyacrylamide containing 8.5 M urea in 0.5× TBE buffer). The gel Plasmid interference assay Plasmid interference assay was performed in *E. coli* strain Arctic Express (DE3) carrying TnpB- and reRNA-encoding plasmids (pGB74-76). Cells were grown at 37° C. to an OD600 of approximately 0.5, then electroporated with 100 ng of target plasmid (pGB72) constructed from pSG4K5 (a gift from Xiao Wang, Addgene Plasmid #74492). After 1 hour, co-transformed cells were further serially diluted by 10-fold dilution, and grown on plates containing IPTG (0.1 mM), gentamicin (10 μg / ml), carbenicillin (100 μg / ml) and kanamycin (50 μg / ml) at 25° C., 30° C. or 37° C. for 16 to 44 hours.
[0167] TnpB-induced DNA cleavage in HEK293T cells HEK293T cells (catalog number CRL-3216) purchased from ATCC were cultured in Dulbecco's Modified Eagle Medium (DMEM) (Gibco) supplemented with 10% fetal bovine serum (Gibco), penicillin (100 U / ml) and streptomycin (100 μg / ml) (Thermo Fisher Scientific). On the day before transfection, cells were seeded at 1.4 × 10 5Cells were seeded in 24-well plates at a cell / well density. The transfection mixture was prepared by mixing 1 μg of plasmid (pRZ122-127) encoding NLS-tagged TnpB and its reRNA with 100 μl of serum-free DMEM and 2 μl of TurboFect transfection reagent (Thermo Fisher Scientific). After incubation at room temperature for 15 minutes, the transfection mixture was added to the cells in droplet form. Transfected cells were grown at 37°C and 5% CO2 for 72 hours.
[0168] Indel Characterization Transfected HEK293T cells were trypsin-treated, and their genomic DNA was extracted using QuickExtract solution (Lucigen). Two rounds of PCR were performed to amplify the DNA regions surrounding each target site and to add sequences necessary for Illumina sequencing and indexing. Briefly, 1–4 μl of DNA lysate was used in the primary PCR with target genomic locus-specific primers, with the 5' ends being Illumina read 1 and read 2 sequences, using Hot Start Phusion polymerase (Thermo Fisher Scientific) to reach a final volume of 20 μl. The thermocycler setting consisted of an initial denaturation at 98°C for 30 seconds, 15 cycles of 98°C for 15 seconds, 56.8°C for 15 seconds, and 72°C for 30 seconds, and a final incubation at 72°C for 5 minutes. The resulting amplification product was purified using 1.8 × volume magnetic beads (Lexogen) and eluted in 30 μl. 6 μl of the elution mixture was used as a template for the second round of PCR, with a final volume of 30 μl. The P5 and P7 adapters required for Illumina sequencing were indexed and added using the Lexogen PCR Add-on Kit (Lexogen) along with the i7 6 nt index set (Lexogen). The thermocycler settings consisted of an initial denaturation at 98°C for 30 seconds, 15 cycles of 98°C for 10 seconds, 65°C for 20 seconds, and 72°C for 30 seconds, and a final incubation at 72°C for 1 minute. To ensure the purity of the PCR product, an additional wash was performed using 0.9 × volume of magnetic beads (Lexogen). Barcoded and purified DNA samples were quantified using a Qubit 4 Fluorometer (Thermo Fisher Scientific), analyzed using a BioAnalyzer (Agilent), pooled in equimolar ratios, and paired-end sequencing (2 × 75 bp) using a MiniSeq High Output Reagent Kit, 150 cycles (Illumina) on a MiniSeq system (Illumina).Insertion or deletion mutations (indels) were analyzed using CRISPResso2 (Clement et al., 2019) with the following parameters: minimum 70% homology for alignment with amplified product sequences, a 10 bp quantitative window, substitutions ignored to avoid false positives, a phred33 score > 10 for average reads, and single base pair quality.
[0169] Example 1: Establishment of the biochemical function of TnpB in the ISDra2 transposition factor of Deinococcus radiodurans. Insertion sequences (IS) are simple, widely found mobile gene elements (MGEs) that contain only genes related to transposition and the regulation of transposition. The IS200 / IS605 family of transposition elements are among the simplest and oldest mobile gene elements (MGEs) (Siguier et al., 2014). Typically, they hold near-terminal palindromic elements (LE and RE) at the MGE terminus and tnpA and tnpB genes in different configurations. However, some MGEs in this family contain standalone tnpA or tnpB genes (ISfinder database) (Siguier et al., 2006). The best experimentally characterized IS608 and IS200 / IS605MGE genes of Helicobacter pylori (Hp) and Deinococcus radiodurans (Dra) ISDra2, respectively, consist of partially duplicated tnpA and tnpB genes that are flanked at the left-end (LE) and right-end (RE) incomplete palindromic sequences (Figure 1A) (Kersulyte et al., 2002; Pasternak et al., 2010). Transposition occurs in conjunction with DNA replication via a "exfoliation and pasting" mechanism involving essential single-stranded DNA intermediates (Hoang et al., 2010).
[0170] The TnpA transposase, encoded by tnpA, is sufficient to facilitate IS mobility both in cellular and in vitro. The TnpA tyrosine Y1 transposase catalyzes both the excision and insertion of ssDNA intermediates. TnpA is a very small (approximately 18 kDa) protein that forms a dimer and contains a complex active site consisting of catalytic tyrosine in one monomer and a metal-binding HUH motif in the other monomer. It cleaves the transposon-coding DNA strand near the "TTAC" (IS608) or "TTGAC" (ISDra2) sequence, which generates a circular single-stranded (ss)DNA intermediate (Figure 1B) (Guynet et al., 2008; Pasternak et al., 2010). The integration reaction specifically occurs within nearby ssDNA of the same sequence, completing the transposition cycle without target site replication (Guynet et al., 2008; Pasternak et al., 2010). Interestingly, target site selection occurs not through direct sequence reading by TnpA, but through base-pairing interactions involving transposon LE factor sequences (Barabas et al., 2008; He et al., 2011). The molecular mechanism of transposition in the IS607 family is not well understood. It may require TnpA serine family transposases and involve a double-stranded (ds)DNA intermediate (Boocock and Rice, 2013; Chen et al., 2018; Kersulyte et al., 2000).
[0171] While the function of TnpA in transposition is well established, the role of TnpB remains unknown. TnpB is thought to be involved in the negative regulation of transposon excision and insertion, although it is not essential for transposition (Kersulyte et al., 2000, 2002; Pasternak et al., 2013). Interestingly, bioinformatics recognition of a conserved RuvC-like active site in the TnpB sequence has led to the speculation that TnpB may be an ancestor of the Cas9 and Cas12 nucleases adopted by the CRISPR-Cas system (Kapitonov et al., 2016; Makarova et al., 2020). However, neither the role of the RuvC motif in transposition nor the nuclease activity of TnpB has been experimentally demonstrated.
[0172] To establish the biochemical function of TnpB in the Deinococcus radiodurans ISDra2 transposition factor, we attempted to isolate and biochemically characterize the TnpB protein. For this purpose, we expressed the E. coli tnpB gene (1227 bp) fused to a sequence encoding a 10×HIs-MBP (maltose-binding protein) purification tag. 2+Initial attempts to purify TnpB from cells extracted by affinity chromatography revealed very low yields of intact TnpB protein (Figure 2A). However, co-expression of tnpB with the complete ISDra2 transposon (with inactivated tnpA) resulted in a significant increase in TnpB yield, suggesting that several transposon factors may contribute to stable TnpB expression (Figures 2B and 2C). Subsequent analysis of TnpB samples revealed that RNA was co-purified with TnpB (Figure 2D). To characterize TnpB-binding RNA, we performed small RNA sequencing (sRNA-seq) and identified enrichment of non-coding RNA (approximately 150 nt) derived from the ISDra2 transposon RE factor, which we named reRNA (Figures 1C and 1D). The reRNA co-purified with TpnB matched the 3' end of the tnpB gene and RE sequence, except for the last approximately 16 nt at the 3' end derived from the plasmid DNA sequence flanking the IS200 / IS605 transposon (Figure 1D). Enrichment of non-coding RNAs that associate with tnpB encoding the IS200 / IS605 family transposons has been previously reported for Halobacterium salinarum (Gomes-Filho et al., 2015). In summary, these data indicate that TnpB forms a ribonucleoprotein (RNP) complex (similar to the Cas9 or Cas12 complex with gRNA) with reRNA derived from the transposon 3' end. In the latter case, the variable sequence portion of the gRNA corresponds to the spacer sequence in the CRISPR array.
[0173] Example 2: RNA that associates with the TnpB protein functions as a guide sequence. We hypothesized that a 3'-terminal reRNA of approximately 16 nt (Figure 1D), derived from DNA adjacent to a transposon and inherently variable, could function as a guide sequence to guide TnpB to its target and activate DNA cleavage via a RuvC-like active site. To test this hypothesis, we employed PAM (protospacer-adjacent motif) recognition assays previously developed for Cas9 and Cas12 nucleases (Karvelis et al., 2015, 2019). Briefly, we first designed reRNA variants in which the 3'-terminal TnpB reRNA sequence derived from the plasmid was replaced with a 16 nt or 20 nt sequence matching the next target in a 7N randomized plasmid library (Figures 3A and 4A). Next, following transformation and expression of E. coli, cell lysates containing the TnpB RNP complex were used directly to establish randomized plasmid library cleavage. The resulting DNA ends were repaired by T4 DNA polymerase, underwent adapter ligation, and were PCR amplified and sequenced. Analysis of adapter-linked fragments revealed enrichment of the product with adapters at target sites of 21–22 bp and 15 bp from the randomized region indicating plasmid library cleavage by the TnpB RNP complex (Figures 3B and 4B). Analysis of the adapter linkage sites on the target (TS) and non-target NTS strands suggested a stepped cleavage region that generates a 5'-overhang. Further analysis of the DNA fragments revealed enrichment of the "TTGAT" sequence in the randomized 7N region 5' upstream of the target sequence. In particular, the TTGAT sequence licensed for TnpB-mediated plasmid library cleavage matched the target site sequence required for TnpA-mediated ISDra2 transposon excision and insertion (Figures 3C, 4C, and 4D) (Islam et al., 2003). Since this sequence is similar to the protospacer adjacent motif (PAM) sequence required for the initiation of DNA cleavage by Cas9 or Cas12 nucleases, we named it the transposon-associated motif (TAM).Next, to confirm the dsDNA cleavage requirements established using plasmid libraries, we purified TnpB RNPs from E. coli and tested their ability to cleave various dsDNA substrates, including target sequences flanked by the 5'-TTGAT TAM sequence (Figures 3D, 8, 9, and 10). The TnpB complex cleaved plasmid DNA (both supercoiled and uncoiled) containing targets flanked by the TAM sequence (Figures 3E, 4C, and 4D). TAM and target sequences matching the reRNA guide sequence were required for plasmid DNA cleavage (Figure 3F). Mutations in conserved residues in the RuvC-like active site impaired cleavage, indicating that RuvC is responsible for dsDNA cleavage (Figure 3E). Finally, run-off sequencing of the cleavage products confirmed a stepped cleavage pattern 15–21 bp from the TAM, resulting in a 5'-overhang (Figures 3G and 8). In summary, these results demonstrate that in vitro TnpB functions as a TAM-dependent RNA-guided dsDNA nuclease.
[0174] Example 3: TnpB allows for in vivo cutting of the donor joint. To test whether TnpB can induce DSBs at the donor joint (Figure 5A) in cells, we monitored the transformation efficiency of recombinant E. coli hosts expressing the TnpB complex with a plasmid containing a target flanking the TAM and a kanamycin (Kn) resistance gene that enables growth on a Kn-supplemented agar plate. Serial dilutions of transformants revealed plasmid interference in cells containing TnpB variants with intact RuvC-like active sites. Plasmid interference was particularly pronounced at lower temperatures (Figures 5B and 11A, 11B). Thus, these results confirm that TnpB is capable of cleaving the donor joint in vivo.
[0175] Example 4: TnpB can mediate targeted genome modification in cells. We investigated whether TnpB could be used for targeted genome modification in human HEK293T cells. A TnpB protein with a nuclear localization sequence (NLS) and a plasmid encoding a reRNA construct targeting human genomic DNA (gDNA) were transiently transfected into HEK293T cells (Figure 7A). After 72 hours, gDNA was extracted and sequenced to analyze the presence of insertions and deletions (indels) at target cleavage sites indicating DSB repair events. In the two tested sites (AGBL1-2 and EMX1-1), TnpB introduced mutations at a frequency of 10–20%, similar to the levels observed in CRISPR-Cas9 and Cas12-based editing (Cong et al., 2013; Jinek et al., 2013; Liu et al., 2019; Mali et al., 2013; Pausch et al., 2020; Zetsche et al., 2015) (Figure 7B). The AGBL1-1 and EMX1-2 sites were moderately modified (1–5%), while no indels were detected in the HPRT1 site. Further analysis of the obtained indels revealed a dominant deletion at the cleavage site (Figure 7C), similar to the mutation profile resulting from Cas12 cleavage (Pausch et al., 2020; Zetsche et al., 2015).
[0176] In summary, these results demonstrate that extremely small RNA-guided TnpB nucleases are capable of cleaving eukaryotic gDNA and can be employed as tools for genome editing, providing a novel class of extremely small non-Cas nucleases with different biochemical requirements for genome editing applications. The table below provides a comparison of RNA-guided TnpB nucleases with Cas9 and Cas12 nucleases. [Table 1] The examples described herein are to be understood as descriptive examples of embodiments of the present invention. Further embodiments and examples are conceivable. Any feature described in relation to any one example or embodiment may be used alone or in combination with other features. Furthermore, any feature described in relation to any one example or embodiment may also be used in combination with one or more features of any other example or embodiment, or in any combination of any other example or embodiment. Furthermore, equivalents and modifications not described herein may also be adopted within the scope of the invention as defined in the claims. [Item 1] A method for cleaving a polynucleotide using an effector complex, wherein the polynucleotide comprises a target sequence, and the effector complex is: (a) Proteins containing or consisting of TnpB protein; and (b) (i) a polynucleotide targeting segment comprising a guide sequence capable of hybridizing to the target sequence; and (ii) A protein-binding segment that enables RNA to bind to the TnpB protein and form the effector complex. The RNA containing The method comprises the step of bringing the polynucleotide into contact with the effector complex, thereby enabling the TnpB protein to cleave the polynucleotide. method. [Item 2] The method according to item 1, wherein the TnpB protein has the amino acid sequence of a protein obtained from the tnpB gene of a mobile genetic element in the IS200 / IS605 family or IS607 family. [Item 3] The method according to item 2, wherein the TnpB protein has the amino acid sequence of a protein obtained from the tnpB gene of a mobile genetic element from the IS200 / IS605 family. [Item 4] The method according to item 3, wherein the TnpB protein has the amino acid sequence of a protein obtained from the tnpB gene of the mobile genetic element ISDra2 from Deinococcus radiodurans. [Item 5] The method according to any one of items 2 to 4, wherein the TnpB protein has at least 85% sequence identity with the amino acid sequence of the protein obtained from the tnpB gene. [Item 6] The method according to any one of items 1 to 5, wherein the TnpB protein comprises or consists of the amino acid sequence of SEQ ID NO: 1, or the TnpB protein comprises or consists of an amino acid sequence having at least 85% sequence identity with SEQ ID NO: 1. [Item 7] The RNA is 50 to 300 nucleotides in length, as described in any one of items 1 to 6. [Item 8] The RNA is 100 to 200 nucleotides in length, as described in item 7. [Item 9] The RNA is 140-150 nucleotides in length, as described in item 8. [Item 10] The method according to any one of items 1 to 9, wherein the guide sequence of the RNA is 10 to 30 nucleotides in length. [Item 11] The method according to item 10, wherein the guide sequence of the RNA is 15 to 25 nucleotides in length. [Item 12] The method according to any one of items 1 to 11, wherein the protein-binding segment of the RNA includes an inverted repeat sequence. [Item 13] The method according to any one of items 1 to 12, wherein the protein-binding segment of the RNA comprises at least a partial palindrome sequence, and the at least partial palindrome sequence is capable of forming a hairpin or an incomplete hairpin. [Item 14] The method according to any one of items 1 to 13, wherein the protein-binding segment of the RNA comprises a sequence from the right-end incomplete palindromic sequence of a mobile gene in the IS200 / IS605 family. [Item 15] The method according to any one of items 1 to 13, wherein the protein-binding segment of the RNA includes a sequence from the rightmost sequence of a mobile gene in the IS607 family. [Item 16] The method according to item 14 or item 15, wherein the protein-binding segment of the RNA includes the sequence from the right end of the insertion sequence containing the gene encoding the TnpB protein. [Item 17] The method according to any one of items 1 to 16, wherein the protein-binding segment comprises a sequence including SEQ ID NO: 2. [Item 18] The method according to any one of items 1 to 17, wherein the protein-binding segment comprises SEQ ID NO: 2 and the TnpB protein comprises SEQ ID NO: 1. [Item 19] The method according to any one of items 1 to 18, wherein the polynucleotide is double-stranded DNA. [Item 20] The polynucleotide is the method described in item 19, wherein the polynucleotide comprises a double-stranded TnpB association sequence motif. [Item 21] The method according to item 20, wherein the double-stranded TnpB association sequence motif is located at 5' of the target sequence that points to the non-target strand of the double-stranded DNA. [Item 22] The method according to item 20 or item 21, wherein the double-stranded TnpB association sequence motif of the polynucleotide is 2 to 6 nucleotides in length. [Item 23] The method according to any one of items 20 to 22, wherein the double-stranded TnpB association sequence motif of the polynucleotide is TTGAT. [Item 24] The method according to item 23, wherein the double-stranded TnpB association sequence motif of the polynucleotide is TTGAT, the TnpB protein comprises or consists of the amino acid sequence of SEQ ID NO: 1, or an amino acid sequence having at least 85% sequence identity with SEQ ID NO: 1, and the protein-binding segment of the RNA comprises SEQ ID NO: 2. [Item 25] The method according to any one of items 1 to 24, wherein the polynucleotide is supercoiled, relaxed, or unwound double-stranded DNA. [Item 26] The method according to any one of items 1 to 25, wherein the polynucleotide is double-stranded DNA, and the break of the double-stranded DNA produces a stepped double-strand break region. [Item 27] The method according to item 26, wherein the stepped double-strand break section has a 5' overhang. [Item 28] The method according to any one of items 1 to 27, wherein the polynucleotide is double-stranded DNA, and the double-stranded DNA is cleaved to produce a blunt-end double-strand break. [Item 29] The method according to any one of items 1 to 28, wherein the polynucleotide containing the target sequence is chromosomal DNA. [Item 30] The method according to any one of items 1 to 28, wherein the polynucleotide containing the target sequence is extrachromosomal DNA. [Item 31] The method according to any one of items 1 to 18, wherein the polynucleotide is single-stranded DNA. [Item 32] The method described above is an in vivo method, as described in any one of items 1 to 31. [Item 33] The method described above is an ex vivo method, as described in any one of items 1 to 31. [Item 34] The method described above is an in vitro method, as described in any one of items 1 to 31. [Item 35] The polynucleotide is present in a cell, as described in any one of items 1 to 34. [Item 36] The cell is a prokaryotic cell, as described in item 35. [Item 37] The cell is a eukaryotic cell, as described in item 35. [Item 38] The method described in item 37, wherein the cells are non-human animal cells. [Item 39] The cell is a human cell, as described in item 37. [Item 40] The method according to any one of items 37 to 39, wherein the cells are stem cells. [Item 41] The method according to item 40, wherein the stem cells are human stem cells, and are not pluripotent stem cells. [Item 42] The method described in item 40, wherein the cells are induced pluripotent stem cells. [Item 43] The method described in item 37, wherein the cells are plant cells. [Item 44] The method according to any one of items 1 to 43, wherein the protein comprises one or more nuclear localization signals on the amino and / or carboxyl terminus of the protein. [Item 45] The method according to any one of items 1 to 44, wherein the contact comprises introducing into a cell: (1) the protein comprising the TnpB protein, or a second polynucleotide sequence comprising a sequence encoding the protein; and (2) the RNA, or a third polynucleotide sequence comprising a sequence encoding the RNA. [Item 46] (1) and (2) are the method of item 45, wherein at least one vector is introduced into the cell. [Item 47] The method according to item 46, wherein the at least one vector comprises a first vector including (1) and a second vector including (2). [Item 48] The method according to item 46 or 47, wherein the at least one vector is at least one nonviral vector, optionally a plasmid, or a nonviral particle such as a liposome or exosome. [Item 49] The method according to item 46 or item 47, wherein the at least one vector is at least one viral vector, optionally a retroviral vector, a lentiviral vector, an adenovirus vector, an adeno-associated virus (AAV) vector, or a herpes simplex virus vector. [Item 50] The method according to item 49, wherein the at least one viral vector is an AAV vector. [Item 51] The method according to any one of items 45 to 50, wherein the contact comprises introducing the second polynucleotide sequence and the third polynucleotide sequence into the cell, wherein the protein comprising the TnpB protein and the RNA are expressed in the cell. [Item 52] The method according to any one of items 45 to 51, wherein the second polynucleotide sequence comprises a first regulator operably ligated to the sequence encoding the protein to control the expression of the protein, and / or the third polynucleotide sequence comprises a second regulator operably ligated to the sequence encoding the RNA to control the expression of the RNA. [Item 53] The method according to any one of items 45 to 50, wherein the protein containing the TnpB protein and the RNA are introduced into the cell in the form of an effector complex. [Item 54] The method according to any one of items 45 to 53, wherein the introduction into the cell is by microinjection or electroporation, and is optionally combined with the use of liposomes. [Item 55] The method according to any one of items 45 to 53, wherein the introduction into the cell is a chemical-based transfection such as lipofection, calcium phosphate transfection, or cationic polymer transfection. [Item 56] The method according to any one of items 1 to 55, wherein the contact occurs under conditions that allow for non-homologous end joining or homologous recombination repair of the cleaved polynucleotide. [Item 57] The method according to item 56, further comprising the step of contacting the polynucleotide with a donor polynucleotide for homologous recombination repair. [Item 58] The method described above is not a method for treating the body of a human or animal, but is a method described in any one of items 1 to 57. [Item 59] The method described in any one of items 1 to 58, wherein the method is not a method for altering human germline genetic identity. [Item 60] The method described above is not the use of a human fetus for an industrial purpose, and is the method described in any one of items 1 to 59. [Item 61] RNA for guiding the effector complex to the target region in a polynucleotide, (i) a polynucleotide targeting segment comprising a guide sequence capable of hybridizing to a target sequence in the target region of the polynucleotide; and (ii) A protein-binding segment that enables the RNA to bind to the TnpB protein and form the effector complex. RNA possessing the following characteristics. [Item 62] The RNA is the RNA described in item 61, having a length of 50 to 300 nucleotides. [Item 63] The RNA in question is the RNA described in item 62, having a length of 100 to 200 nucleotides. [Item 64] The RNA in question is the RNA described in item 63, having a length of 140 to 150 nucleotides. [Item 65] The guide sequence is an RNA sequence described in any one of items 61 to 64, having a length of 10 to 30 nucleotides. [Item 66] The aforementioned guide sequence is the RNA described in item 65, having a length of 15 to 25 nucleotides. [Item 67] The protein-binding segment comprises an inverted repeat sequence, and is an RNA as described in any one of items 61 to 66. [Item 68] The RNA according to any one of items 61 to 67, wherein the protein-binding segment comprises at least a partial palindromic sequence, and the at least partial palindromic sequence is capable of forming a hairpin or an incomplete hairpin. [Item 69] The protein-binding segment comprises the RNA described in any one of items 61 to 68, which includes a sequence from a right-end incomplete palindrome sequence of a mobile gene in the IS200 / IS605 family, or a sequence from the right end of a mobile gene in the IS607 family. [Item 70] The RNA according to item 69, wherein the protein-binding segment includes the sequence from the right end of the insertion sequence containing the gene encoding the TnpB protein. [Item 71] The protein-binding segment is RNA as described in any one of items 61 to 70, including Sequence ID: 2. [Item 72] The TnpB protein to which the protein-binding segment can bind is the RNA according to any one of items 61 to 71, having the amino acid sequence of a protein obtained from the tnpB gene of a mobile genetic element in the IS200 / IS605 family or the IS607 family. [Item 73] The TnpB protein is the RNA described in item 72, having the amino acid sequence of a protein obtained from the tnpB gene of the IS200 / IS605 family of mobile genetic elements. [Item 74] The TnpB protein is the RNA described in item 73, having the amino acid sequence of a protein obtained from the tnpB gene of the mobile genetic element ISDra2 from Deinococcus radiodurans. [Item 75] The TnpB protein to which the protein-binding segment can bind comprises or consists of the amino acid sequence of SEQ ID NO: 1, or comprises or consists of an amino acid sequence having at least 85% sequence identity with SEQ ID NO: 1, as described in any one of items 61 to 74. [Item 76] The protein-binding segment comprises SEQ ID NO: 2, and the TnpB protein to which the protein-binding segment can bind comprises SEQ ID NO: 1, as described in any one of items 61 to 75. [Item 77] The RNA according to any one of items 61 to 76, wherein the polynucleotide containing the target sequence, which is hybridizable to the guide sequence, is DNA. [Item 78] The RNA according to any one of items 61 to 76, wherein the polynucleotide containing the target sequence, which is hybridizable to the guide sequence, is RNA. [Item 79] An isolated RNA, as described in any one of items 61 to 78. [Item 80] The RNA is a designed, non-naturally occurring RNA, as described in any one of items 61 to 79. [Item 81] An effector complex for binding to a target region in a polynucleotide, wherein the effector complex comprises a protein and RNA, wherein the protein comprises or consists of a TnpB protein, and the RNA is: (i) a polynucleotide targeting segment comprising a guide sequence capable of hybridizing to a target sequence comprised in said target region; and (ii) a protein binding segment that binds to said TnpB protein An effector complex comprising [Item 82] The effector complex according to Item 81, wherein said TnpB protein has the amino acid sequence of a protein obtained from the tnpB gene of a mobile genetic element in the IS200 / IS605 family or the IS607 family, or is a variant having an amino acid sequence that has at least 85% sequence identity therewith [Item 83] The effector complex according to Item 82, wherein said TnpB protein has the amino acid sequence of a protein obtained from the tnpB gene of a mobile genetic element in said IS200 / IS605 family, or is a variant having an amino acid sequence that has at least 85% sequence identity therewith [Item 84] The effector complex according to Item 83, wherein said TnpB protein has the amino acid sequence of a protein obtained from the tnpB gene of the mobile genetic element ISDra2 from Deinococcus radiodurans, or is a variant thereof having an amino acid sequence that has at least 85% sequence identity therewith [Item 85] The effector complex according to any one of Items 81 to 84, wherein said TnpB protein comprises or consists of the amino acid sequence of SEQ ID NO: 1, or said TnpB protein comprises or consists of an amino acid sequence having at least 85% sequence identity with SEQ ID NO: 1 [Item 86] The effector complex according to any one of Items 81 to 85, wherein said RNA is as defined in any one of Items 61 to 80 [Item 87] The effector complex according to any one of Items 81 to 86, wherein said protein binding segment comprises SEQ ID NO: 2 [Item 88] The effector complex according to any one of Items 81 to 87, wherein the protein comprises one or more nuclear localization signals at the amino or carboxyl terminus of the protein, and / or the protein comprises one or more cell membrane penetrating peptides at the amino or carboxyl terminus of the protein. [Item 89] The effector complex according to any one of Items 81 to 88, wherein the effector complex is for modifying the target region of the polynucleotide. [Item 90] The effector complex according to any one of Items 81 to 89, wherein the effector complex is for cleaving the target region of the polynucleotide using the TnpB protein. [Item 91] The effector complex according to any one of Items 81 to 89, wherein the effector complex comprises a TnpB protein having an inactivated nuclease domain. [Item 92] The effector complex according to any one of Items 81 to 91, comprising one or more effector molecules for modifying the target region. [Item 93] The effector complex according to Item 92, wherein the one or more effector molecules are selected from TnpB protein, ribonuclease, nickase, base editor, epigenetic modifier, transposase, recombinase, and endonuclease that is not reverse transcriptase. [Item 94] The effector complex according to Item 93, wherein the base editor is a deaminase, and optionally is a cytidine deaminase and / or an adenosine deaminase. [Item 95] The effector complex according to Item 94, wherein the effector complex comprises a cytidine deaminase and a uracil glycosylase inhibitor. [Item 96] An effector complex according to any one of items 81 to 92, comprising one or more effector molecules for labeling the target region. [Item 97] The effector complex described in item 96, wherein the label is a fluorescent protein. [Item 98] An effector complex according to any one of items 81 to 91, comprising one or more effector molecules for increasing or decreasing the transcription or translation of the target region. [Item 99] The effector complex according to item 98, wherein one or more of the effector molecules are one or more transcription activators. [Item 100] The effector complex of item 98, wherein one or more of the effector molecules are one or more transcriptional repressors. [Item 101] The effector complex according to any one of items 92 to 100, wherein the protein is a fusion protein comprising the TnpB protein and the one or more effector molecules. [Item 102] An effector complex according to any one of items 81 to 101, bonded to a solid support. [Item 103] The effector complex described in any one of items 81 to 102, wherein the polynucleotide is double-stranded DNA. [Item 104] The double-stranded DNA is supercoiled, relaxed, or unwound, as described in item 103 of the effector complex. [Item 105] The effector complex according to item 103 or item 104, wherein the polynucleotide includes a TnpB association sequence motif located at 5' of the target sequence, which refers to the non-target strand of the double-stranded DNA. [Item 106] The effector complex according to item 105, wherein the TnpB association sequence motif of the polynucleotide is 2 to 6 nucleotides in length. [Item 107] The effector complex according to item 105 or item 106, wherein the TnpB association sequence motif of the polynucleotide is TTGAT. [Item 108] The effector complex according to any one of items 103 to 107, wherein the TnpB protein is capable of cleaving the double-stranded DNA to produce staggered double-strand breaks. [Item 109] The aforementioned stepped double-strand break section has a 5' overhang, as described in item 108. [Item 110] The effector complex according to any one of items 103 to 107, wherein the TnpB protein is capable of cleaving the double-stranded DNA to produce blunt-end double-strand breaks. [Item 111] The effector complex described in any one of items 81 to 102, wherein the polynucleotide is single-stranded DNA. [Item 112] The effector complex described in any one of items 81 to 102, wherein the polynucleotide is RNA. [Item 113] An isolated effects unit, as described in any one of items 81 through 112. [Item 114] The aforementioned effector complex is a designed, non-naturally occurring effector complex, as described in any one of items 81 to 112. [Item 115] A fusion protein for forming an effector complex having RNA as described in any one of items 61 to 80, wherein the fusion protein comprises a TnpB protein and (i) one or more nuclear localization signals and / or cell membrane permeable peptides on the amino or carboxyl terminus of the fusion protein and / or (ii) one or more effector molecules. [Item 116] The fusion protein according to item 115, wherein said TnpB protein has an amino acid sequence of a protein obtained from the tnpB gene of a mobile genetic element in the IS200 / IS605 family or the IS607 family, or is a variant having an amino acid sequence having at least 85% sequence identity to said sequence. [Item 117] The fusion protein according to item 116, wherein said TnpB protein has an amino acid sequence of a protein obtained from the tnpB gene of a mobile genetic element from said IS200 / IS605 family, or is a variant thereof having an amino acid sequence having at least 85% sequence identity to said sequence. [Item 118] The fusion protein according to item 117, wherein said TnpB protein has an amino acid sequence of a protein obtained from the tnpB gene of the mobile genetic element ISDra2 from *Deinococcus radiodurans*, or is a variant thereof having an amino acid sequence having at least 85% sequence identity to said sequence. [Item 119] The fusion protein according to any one of items 115 to 118, wherein said TnpB protein comprises or consists of the amino acid sequence of SEQ ID NO: 1, or comprises or consists of an amino acid sequence having at least 85% sequence identity with SEQ ID NO: 1. [Item 120] The fusion protein according to any one of items 115 to 119, wherein said one or more effector molecules are for modifying, labeling a target region in a polynucleotide, or increasing or decreasing the transcription or translation of said target region. [Item 121] The fusion protein according to any one of items 115 to 120, wherein said one or more effector molecules are one or more selected from the group consisting of said TnpB protein, ribonuclease, nickase, base editor, epigenetic modifier, transposase, recombinase, reverse transcriptase, label, transcriptional activator, and endonuclease that is not a transcriptional repressor. [Item 122] The base editor is a deaminase, optionally cytidine deaminase and / or adenosine deaminase, as described in item 121, which is a fusion protein. [Item 123] A fusion protein as described in item 122, comprising cytidine deaminase and uracil glycosylase inhibitor. [Item 124] The label is a fluorescent protein or a reporter enzyme, as described in item 121, for the fusion protein. [Item 125] A mutant TnpB protein comprising a mutation, wherein the mutation is for inactivating the nuclease domain of the protein, the mutant TnpB protein is configured to bind to an RNA described in any one of items 61 to 80, and optionally, the mutant TnpB protein is the TnpB protein of a fusion protein described in any one of items 115 to 124. [Item 126] DNA encoding an RNA as described in any one of items 61 through 80. [Item 127] DNA as described in item 126, further encoding a protein comprising the TnpB protein as defined in any one of items 82 to 85, or further encoding a protein comprising the mutant TnpB protein of item 125. [Item 128] DNA as described in item 126, further encoding a fusion protein as described in any one of items 115 to 124. [Item 129] DNA encoding a fusion protein as described in any one of items 115 to 127. [Item 130] DNA encoding the mutant TnpB protein described in item 125. [Item 131] The DNA is a designed, non-naturally occurring DNA, as described in any one of items 126 to 130. [Item 132] A recombinant expression vector comprising DNA as described in any one of items 126 to 131, optionally being a plasmid or a viral vector such as an AAV vector. [Item 133] A host cell containing DNA as described in any one of items 126 to 131, or a recombinant expression vector as described in any one of item 132. [Item 134] A composition comprising RNA as described in any one of items 61 to 80, an effector complex as described in any one of items 81 to 114, a fusion protein as described in any one of items 115 to 124, a mutant TnpB protein as described in item 125, DNA as described in any one of items 126 to 131, a recombinant expression vector as described in item 132, or a host cell as described in item 133, and a buffer. [Item 135] The buffer is a pharmaceutically acceptable buffer, as described in item 134. [Item 136] An in vitro or ex vivo method for producing RNA according to any one of items 61 to 80, comprising the steps of expressing DNA according to any one of items 126 to 128, or chemically synthesizing the RNA. [Item 137] An in vitro or ex vivo method for producing a fusion protein according to any one of items 115 to 124, or a mutant TnpB protein according to item 125, comprising the step of expressing DNA according to item 129 or 130. [Item 138] An in vitro or ex vivo method for producing an effector complex according to any one of items 81 to 114, comprising the steps of: contacting the RNA with the protein to form the effector complex; optionally, expressing DNA encoding the protein, including the TnpB protein; and expressing DNA encoding the RNA. [Item 139] The method described above is the method described in any one of items 136 to 138, performed in vitro in cells. [Item 140] The method described above is the method described in any one of items 136 to 138, performed in vitro in a cell-free system. [Item 141] The method according to any one of items 136 to 140, comprising the step of purifying the produced RNA, the produced fusion protein, the produced mutant TnpB protein, or the produced effector complex. [Item 142] A system for modifying a target region in a polynucleotide, wherein the target region includes a target sequence, and the system is: a) A protein containing or consisting of the TnpB protein, or DNA or RNA encoding the said protein, and b) RNA, or DNA encoding the RNA. The RNA is: (i) a polynucleotide targeting segment comprising a sequence complementary to the target sequence; and (ii) Protein-binding segment that binds to the TnpB protein A system that includes this. [Item 143] (a) the protein and (b) the RNA, the system according to item 142. [Item 144] The system according to item 142, comprising (a) the DNA encoding the protein, and (b) the DNA encoding the RNA. [Item 145] (a) the RNA encoding the protein, and (b) the system according to item 142, comprising the RNA. [Item 146] The system according to item 142, comprising (a) the protein and (b) the DNA encoding the RNA. [Item 147] The RNA is a system as defined in any one of items 61 to 80, or as described in any one of items 142 to 146. [Item 148] The protein is a system as defined in any one of items 82 to 85, or as described in any one of items 142 to 147. [Item 149] The system according to any one of items 142 to 147, wherein the protein is a fusion protein as defined in any one of items 115 to 124. [Item 150] The TnpB protein comprises an inactivated nuclease domain, as described in any one of items 142 to 149. [Item 151] The polynucleotide is a system as defined in any one of items 142 to 150, as defined in any one of items 19 to 23 or 29 to 31. [Item 152] (a) and / or (b) are included in at least one vector, the system described in any one of items 142 to 151. [Item 153] The system according to item 152, wherein at least one of the vectors is at least one nonviral vector. [Item 154] The system according to item 153, wherein at least one of the nonviral vectors is at least one plasmid, or at least one nonviral particle such as a liposome or exosome. [Item 155] The system according to item 152, wherein at least one of the vectors is at least one viral vector. [Item 156] The system according to item 155, wherein at least one of the viral vectors is selected from a retroviral vector, a lentiviral vector, an adenovirus vector, an adeno-associated virus (AAV) vector, or a herpes simplex virus vector. [Item 157] The system according to item 156, wherein at least one of the viral vectors is at least one AAV vector. [Item 158] The aforementioned system is a designed, non-naturally occurring system, as described in any one of items 142 to 157. [Item 159] An effector complex as described in any one of items 81 to 114, or a system as described in any one of items 142 to 158, for use as a drug or for use in diagnosis. [Item 160] Use of an effector complex as described in any one of items 81 to 114, or a system as described in any one of items 142 to 158, for the manufacture of a drug for treating or preventing a disease in a subject, or for the manufacture of a diagnostic composition for diagnosing a disease in a subject. [Item 161] The RNA is the RNA described in any one of items 61 to 80, used in combination with the protein as defined in any one of items 82 to 85 or 125, or with the DNA or RNA encoding the protein, for use as a drug or for use in the diagnosis of a subject. [Item 162] The RNA is the RNA described in any one of items 61 to 80, used in combination with the fusion protein as defined in any one of items 115 to 124, or with the DNA or RNA encoding the fusion protein, for use as a drug or for use in the diagnosis of a subject. [Item 163] A protein as defined in any one of items 82 to 85 or 125, or RNA-encoding DNA as described in item 126, used in conjunction with DNA encoding such protein, for use as a drug or for use in the diagnosis of a subject. [Item 164] A fusion protein as defined in any one of items 115 to 124, or a DNA encoding the RNA described in item 126, used in conjunction with the DNA encoding the fusion protein, for use as a drug or for use in the diagnosis of a subject. [Item 165] DNA or RNA encoding a protein for use as a drug or for use in the diagnosis of a subject, wherein the protein is defined in any one of items 82 to 85 or 125, and the DNA is RNA as described in any one of items 61 to 80, or DNA or RNA used in conjunction with the DNA encoding the RNA. [Item 166] DNA or RNA encoding a fusion protein for use as a drug or for use in the diagnosis of a subject, wherein the fusion protein is defined in any one of items 115 to 124, and the DNA is RNA as described in any one of items 61 to 80, or DNA or RNA used in conjunction with the DNA encoding the RNA. [Item 167] Use of an effector complex described in any one of items 81 to 114, or a system described in any one of items 142 to 158, in an in vitro or ex vivo method for determining the presence of a polynucleotide containing a target sequence in a sample. [Item 168] Use of an effector complex described in any one of items 81 to 114, or a system described in any one of items 142 to 158, in an in vitro or ex vivo method for modifying a target region of a polynucleotide containing a target sequence. [Item 169] Use of an effector complex described in any one of items 81 to 114, or a system described in any one of items 142 to 158, in an in vitro or ex vivo method for genetically modifying cells. [Item 170] The aforementioned cells are used as defined in any one of items 36 to 43, as described in item 169. [Item 171] Genetically modified cells used as a drug in a subject, wherein the genetically modified cells are obtained by a method comprising the step of genetically modifying cells obtained from the subject using a system described in any one of items 142 to 158 or an effector complex described in any one of items 81 to 114. [Item 172] The method described above comprises the step of culturing and increasing the cells obtained from the subject, as described in item 171. [Item 173] The cells obtained from the subject are hematopoietic stem cells and progenitor cells (HSPCs) or T cells, as described in items 171 and 172, which are genetically modified cells. [Item 174] A method for modifying, labeling, or controlling the expression of a target region in a polynucleotide using an effector complex, wherein the target region includes a target sequence. The effector complex is: (i) an effector complex of any of items 81 to 114; (ii) a fusion protein of any of items 115 to 124 and RNA of any of items 61 to 80; or (iii) a mutant TnpB protein of item 125, or a fusion protein containing the mutant TnpB protein and RNA of any of items 61 to 80. The method comprises the step of contacting the polynucleotide with the effector complex, wherein the guide sequence of the RNA hybridizes to the target sequence, and the effector complex modifies or labels the target region, or controls the expression from the target region. method. [Item 175] The method according to item 174, wherein the effector complex comprises one or more effector molecules. [Item 176] The method according to item 175, wherein one or more of the effector molecules are selected from one or more endonucleases that are not mutant TnpB proteins, ribonucleases, nickases, base editors, epigenetic modifiers, transposases, recombinases, reverse transcriptases, labels, transcription activators, and transcription repressors. [Item 177] The method according to any one of items 174 to 176, wherein the effector complex comprises a fusion protein containing the mutant TnpB protein and one or more effector molecules. [Item 178] The method according to any one of items 174 to 177, wherein the polynucleotide is double-stranded DNA. [Item 179] The method according to any one of items 174 to 177, wherein the polynucleotide is single-stranded DNA. [Item 180] The method according to any one of items 174 to 177, wherein the polynucleotide is RNA. [Item 181] The method described above is an in vivo method, as described in any one of items 174 to 180. [Item 182] The method described above is an ex vivo method, as described in any one of items 174 to 180. [Item 183] The method described above is an in vitro method, as described in any one of items 174 to 180. [Item 184] The polynucleotide is present in cells, as described in any one of items 174 to 183. [Item 185] The cell is a prokaryotic cell, as described in item 184. [Item 186] The cell is a eukaryotic cell, as described in item 184. [Item 187] The cells are non-human animal cells, as described in item 186. [Item 188] The cell is a human cell, as described in item 186. [Item 189] The cell is a stem cell, as described in any one of items 186 to 188. [Item 190] The method described in item 189, where the stem cells are human stem cells and are not pluripotent stem cells. [Item 191] The cells are induced pluripotent stem cells, as described in item 189. [Item 192] The cell is a plant cell, as described in item 186. [Item 193] The method according to any one of items 174 to 192, comprising the step of introducing an effector complex into a cell, or introducing a system according to any one of items 142 to 158 into a cell. [Item 194] The method according to item 193, wherein the introduction step comprises electroporation or microinjection, and optionally combined with the use of liposomes. [Item 195] The method according to item 193, wherein the introduction into the cells is a chemical-based transfection such as lipofection, calcium phosphate transfection, or cationic polymer transfection. [Item 196] The introduction into the cell is by delivery using a viral vector, as described in item 193. [Item 197] The method according to item 196, wherein the viral vector is a lentiviral vector, a retroviral vector, or an AAV vector. [Item 198] The method described above is not a method for treating the body of a human or animal, but is a method described in any one of items 174 to 197. [Item 199] The method described above is not a method for altering human germline genetic identity, but is a method described in any one of items 174 to 198. [Item 200] The method described above is the method described in any one of items 174 to 199, not the use of a human fetus for an industrial purpose.
[0177] All publications referred to herein are incorporated by reference to the same extent as each individual publication is specifically and individually shown to be incorporated by reference in whole. References References 1. Anzalone et al., (2020) Genome editing with CRISPR-Cas nucleases, base editors, transposases and prime editors. Nature Biotechnology 38, 824-844 Reference 2. Barabas, O., Ronning, DR, Guynet, C., Hickman, AB, Ton-Hoang, B., Chandler, M., and Dyda, F. (2008). Mechanism of IS200 / IS605 Family DNA Transposases: Activation and Transposon-Directed Target Site Selection. Cell 132, 208-220. Reference 3. Boocock, MR, and Rice, PA (2013). A proposed mechanism for IS607-family serine transposases. Mob DNA 4, 24. Enclosure 4. Chen , W. , Mandali , S. , Hancock , SP , Kumar , P. , Collazo , M. , Cascio , D. , and Johnson , RC (2018). Multiple serine transposase dimers assemble the transposon-end synaptic complex during IS607-family transposition. ELife 7 , e39611. Enclosure 5. Clement , K , Rees , H , Canver , MC , Gehrke , JM , Farouni , R , Hsu , JY , Cole , MA , Liu , DR , Joung , JK , Bauer , DE , et al. (2019). CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nature Biotechnology 37, 224–226. Enclosure 6. Cong , L. , Ran , FA , Cox , D. , Lin , S. , Barretto , R. , Habib , N. , Hsu , PD , Wu , X. , Jiang , W. , Marraffini , LA , et al. (2013). Multiplex Genome Engineering Using CRISPR / Cas Systems. Science 339, 819–823. Enclosure 7. Crooks , GE ( 2004 ). WebLogo: A Sequence Logo Generator. Genome Research 14, 1188–1190. Enclosure 8. Gillmore et al., (2021) CRISPR-Cas9 in vivo gene-editing for transthyretin amyloidosis. NEJM DOI: 10.1056 / NEJMoa2107454, 26 June Reference 9. Gomes-Filho, J.V., Zaramela, L.S., Italiani, V.C. da S., Baliga, N.S., Vencio, R.Z.N., and Koide, T. (2015). Sense overlapping transcripts in IS1341-type transposase genes are functional non-coding RNAs in archaea. RNA Biol 12, 490-500. Reference 10. Guynet, C., Hickman, A.B., Barabas, O., Dyda, F., Chandler, M., and Ton-Hoang, B. (2008). In Vitro Reconstitution of a Single-Stranded Transposition Mechanism of IS608. Molecular Cell 29, 302-312. Reference 11. Hajian et al., (2019) Detection of unamplified target genes via CRISPR-Cas9 immobilized on a graphene field-effect transistor. Nature Biomedical Engineering 3, 427-437 Reference 12. He, S., Hickman, A.B., Dyda, F., Johnson, N.P., Chandler, M., and Ton-Hoang, B. (2011). Reconstitution of a functional IS608 single-strand transpososome: role of non-canonical base pairing. Nucleic Acids Research 39, 8503-8512. Reference 13. Hickman, AB, Chandler, M., Dyda, F. (2010) Integrating prokaryotes and eukaryotes: DNA transposases in light of structure. Crit. Rev. Biochem. Mol. Biol. 45, 50-56. Reference 14. Hoang, BT, Pasternak, C., Siguier, P., Guynet, C., Hickman, AB, Dyda, F., Sommer, S., and Chandler, M. (2010). Single-stranded DNA transposition is coupled to host replication. Cell 142, 398-408. Reference 15. Islam, MS, Hua, Y., Ohba, H., Satoh, K., Kikuchi, M., Yanagisawa, T., and Narumi, I. (2003). Characterization and distribution of IS8301 in the radioresistant bacterium Deinococcus radiodurans. Genes Genet. Syst. 78, 319-327. Reference 16. Jiang, F., and Doudna, J. (2017). CRISPR-Cas9 Structures and Mechanisms. Ann. Rev. Biophys. 46, 505-29. Reference 17. Jinek, M., East, A., Cheng, A., Lin, S., Ma, E., and Doudna, J. (2013). RNA-programmed genome editing in human cells. ELife 2, e00471. Reference 18. Kapitonov, V.V., Makarova, K.S., and Koonin, E.V. (2016). ISC, a Novel Group of Bacterial and Archaeal DNA Transposons That Encode Cas9 Homologs. J. Bacteriol. 198, 797-807. Reference 19. Karvelis, T., Gasiunas, G., Young, J., Bigelyte, G., Silanskas, A., Cigan, M., and Siksnys, V. (2015). Rapid characterization of CRISPR-Cas9 protospacer adjacent motif sequence elements. Genome Biol 16, 253. Reference 20. Karvelis, T., Young, J.K., and Siksnys, V. (2019). A pipeline for characterization of novel Cas9 orthologs. In Methods in Enzymology, (Elsevier), pp. 219-240. Reference 21. Karvelis, T., Bigelyte, G., Young, J.K., Hou, Z., Zedaveinyte, R., Budre, K., Paulraj, S., Djukanovic, V., Gasior, S., Silanskas, A., et al. (2020). PAM recognition by miniature CRISPR-Cas12f nucleases triggers programmable double-stranded DNA target cleavage. Nucleic Acids Res 48, 5016-5023. Entertainment22. Kersulyte , D. , Mukhopadhyay , AK , Shirai , M. , Nakazawa , T. , and Berg , DE (2000). Functional Organization and Insertion Specificity of IS607, a Chimeric Element of Helicobacter pylori. Journal of Bacteriology 182, 5300–5308. Entertainment23. Kersulyte , D. , Velapatino , B. , Dailide , G. , Mukhopadhyay , AK , Ito , Y. , Cahuayme , L. , Parkinson , AJ , Gilman , RH , and Berg , DE (2002). Transposable Element ISHp608 of Helicobacter pylori: Nonrandom Geographic Distribution, Functional Organization, and Insertion Specificity. Journal of Bacteriology 184, 992–1002. Entertainment 24. Knott et al., (2018) CRISPR-Cas guides the future of genetic engineering. Science 361, 866–869 Entertainment 25. Krupovic , M. , Makarova , KS , Forterre , P. , Prangishvili , D. , and Koonin , EV (2014). Casposons: a new superfamily of self-synthesizing DNA transposons at the origin of prokaryotic CRISPR-Cas immunity. BMC Biology 12 , 36 . Reference 26. Li, H., and Durbin, R. (2009). Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25, 1754-1760. Reference 27. Li, H., Handsaker, B., Wysoker, A., Fennell, T., Ruan, J., Homer, N., Marth, G., Abecasis, G., Durbin, R., and 1000 Genome Project Data Processing Subgroup (2009). The Sequence Alignment / Map format and SAMtools. Bioinformatics 25, 2078-2079. Reference 28. Liu, J.-J., Orlova, N., Oakes, B.L., Ma, E., Spinner, H.B., Baney, K.L.M., Chuck, J., Tan, D., Knott, G.J., Harrington, L.B., et al. (2019). CasX enzymes comprise a distinct family of RNA-guided genome editors. Nature 566, 218-223. Reference 29. Ma H, Tu LC, Naseri A, Chung YC, Grunwald D, Zhang S, Pederson T.. (2018) CRISPR-Sirius: RNA scaffolds for signal amplification in genome imaging. Nat Methods. Nov;15(11):928-931. Enclosure 30. [ PubMed ] Ma H, Tu LC, Naseri A, Huisman M, Zhang S, Grunwald D, Pederson T. (2016) Multiplexed labeling of genomic loci with dCas9 and engineered sgRNAs using CRISPRainbow. Nat Biotechnol.May;34(5):528-30. Environmental Protection31. Madeira , F. , Park , YM , Lee , J. , Buso , N. , Gur , T. , Madhusoodanan , N. , Basutkar , P. , Tivey , ARN , Potter , SC , Finn , RD , et al. (2019). Nucleic Acids Res 47, W636-W641. Environmental Protection32. Makarova , KS , Wolf , YI , Iranzo , J , Shmakov , SA , Alkhnbashi , OS , Brouns , SJJ , Charpentier , E , Cheng , D , Haft , DH , Horvath , P , et al. (2020). Evolutionary classification of CRISPR-Cas systems: a burst of class 2 and derived variants. Nat Rev Microbiol 18, 67–83. Environmental Protection33. Mali , P. , Yang , L. , Esvelt , KM , Aach , J. , Guell , M. , DiCarlo , JE , Norville , JE , and Church , GM (2013). RNA-Guided Human Genome Engineering via Cas9. Science 339, 823–826. Reference 34. Maresca et al., (2013). Obligate Ligation-Gated Recombination (ObLiFaRe): Custom-designed nuclease-mediated targeted integration through nonhomologous end joining. Genome Research 23, 539-546. Reference 35. Martin, M. (2011). Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet.Journal 17, 10-12. Reference 36. Mir et al., Heavily and fully modified RNAs guide efficient SpyCas9-mediated genome editing - Nature Communications, 9, Article No. 2641 (2018). Reference 37. Pasternak, C., Ton-Hoang, B., Coste, G., Bailone, A., Chandler, M., and Sommer, S. (2010). Irradiation-Induced Deinococcus radiodurans Genome Fragmentation Triggers Transposition of a Single Resident Insertion Sequence. PLoS Genet 6, e1000799. Entertainment 38. Pasternak , C. , Dulermo , R. , Ton‐Hoang , B. , Debuchy , R. , Siguier , P. , Coste , G. , Chandler , M. , and Sommer , S. (2013). ISDra2 transposition in Deinococcus radiodurans is downregulated by TnpB. Molecular Microbiology 88, 443–455. Environmental Protection39. Pausch , P. , Al-Shayeb , B. , Bisom-Rapp , E. , Tsuchida , CA , Li , Z. , Cress , BF , Knott , GJ , Jacobsen , SE , Banfield , JF , and Doudna , JA (2020). CRISPR-CasΦ from large phages is a hypercompact genome editor. Science 369, 333–337. Enclosure 40 . Robinson , JT , Thorvaldsdottir , H , Winckler , W , Guttman , M , Lander , ES , Getz , G , and Mesirov , JP (2011). Integrative genomics viewer. Nature Biotechnology 29, 24-26. Environmental Protection41. Sajwan S, Mannervik M (2019) Gene activation by dCas9-CBP and the SAM system differ in target preference. Scientific Reports 9: Article no. 18104 Environmental Protection42. Shmakov , S. , Smargon , A. , Scott , D. , Cox , D. , Pyzocha , N. , Yan , W. , Abudayyeh , OO , Gootenberg , JS , Makarova , KS , Wolf , YI , et al. (2017). Diversity and evolution of class 2 CRISPR-Cas systems. Nat Rev Microbiol 15, 169–182. Environmental Protection43. Siguier , P. , Perochon , J. , Lestrade , L. , Mahillon , J. , and Chandler , M. (2006). ISfinder: the reference center for bacterial insertion sequences. Nucleic Acids Res 34, D32–36. Enclosure 44 . http: / / dx.doi.org / 10.1037 / 0021-843X.113.2.202 Siguier, P., Gourbeyre, E., & Chandler, M. (2014). Bacterial insertion sequences: their genomic impact and diversity. FEMS Microbiology Reviews 38, 865–891. Environmental 45 . Takeda , SN , Nakagawa , R. , Okazaki , S. , Hirano , H. , Kobayashi , K. , Kusakizako , T. , Nishizawa , T. , Yamashita , K. , Nishimasu , H. , and Nureki , O. (2020). Structure of the miniature type VF CRISPR-Cas effector enzyme. Molecular Cell S1097276520308352. Environmental Protection46. Xiao , R. , Li , Z. , Wang , S. , Han , R. , and Chang , L. (2021). Structural basis for substrate recognition and cleavage by the dimerization-dependent CRISPR-Cas12f nuclease. Nucleic Acids Research. Environmental Protection47. Xu et al., (2019) Viral Delivery Systems for CRISPR. Viruses, 11, no. 28 Entertainment 48. Zetsche , B , Gootenberg , JS , Abudayyeh , OO , Slaymaker , IM , Makarova , KS , Essletzbichler , P , Volz , SE , Joung , J , van der Oost , J , Regev , A , et al. (2015). Cpf1 Is a Single RNA-Guided Endonuclease of a Class 2 CRISPR-Cas System. Cell 163, 759–771.
Claims
[Claim 1] A method for cleaving a polynucleotide using an effector complex, wherein the polynucleotide comprises a target sequence, and the effector complex is: (a) Proteins containing or consisting of TnpB protein; and (b) (i) a polynucleotide targeting segment comprising a guide sequence capable of hybridizing to the target sequence; and (ii) Protein-binding segment that enables RNA to bind to the TnpB protein and form the effector complex. The RNA containing The method comprises the step of bringing the polynucleotide into contact with the effector complex, thereby enabling the TnpB protein to cleave the polynucleotide. method.