TnpB-like RNA guide nuclease complex

A novel TnpB-like RNA guide nuclease complex enables efficient genome editing by cleaving target DNA at a TAM sequence, overcoming the limitations of PAM sequence dependency in existing technologies, thereby improving the precision and flexibility of genome editing.

JP2026082990APending Publication Date: 2026-05-19NATIONAL INSTITUTE OF ADVANCED INDUSTRIAL SCIENCE & TECHNOLOGY +2
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
NATIONAL INSTITUTE OF ADVANCED INDUSTRIAL SCIENCE & TECHNOLOGY
Filing Date
2026-02-05
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing genome editing technologies, such as CRISPR/Cas9 and CRISPR/Cas12a, are limited by the requirement for a specific PAM sequence, which can vary depending on the type of Cas protein, restricting the flexibility and efficiency of target DNA editing.

Method used

Development of a novel TnpB-like RNA guide nuclease complex that does not rely on the PAM sequence, utilizing a TnpB-like RNA guide nuclease with a specific amino acid sequence and a guide RNA to cleave target DNA at a predetermined site, including a TAM sequence, enabling sequence-specific knockout and knock-in.

Benefits of technology

The TnpB-like RNA guide nuclease complex allows for efficient and flexible genome editing by cleaving target nucleic acids without the need for a PAM sequence, enhancing the precision and applicability of genome editing techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026082990000001_ABST
    Figure 2026082990000001_ABST
Patent Text Reader

Abstract

The objective is to identify and utilize novel RNA guide nucleases applicable to genome editing technology. [Solution] We have discovered a TnpB-like RNA guide nuclease that functions as an RNA guide endonuclease complex applicable to genome editing technology. This provides a TnpB-like RNA guide nuclease complex comprising the TnpB-like RNA guide nuclease and a gRNA containing a guide sequence and a gRNA scaffold sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a novel TnpB-like RNA-guided nuclease complex, a technique for cleaving target DNA using the complex, a genome editing technique associated with the cleavage, and a transcriptional regulation technique in target DNA using the TnpB-like RNA-guided nuclease complex.

Background Art

[0002] Genome editing is a technique for introducing (editing) mutations at targeted locations in the genome of a target organism. The genome editing technique using CRISPR-Cas is based on the discovery of an adaptive immune system in bacteria and archaea against foreign viruses and plasmids, and is a technique that applies Cas, which is an RNA-guided endonuclease involved in such an adaptive immune system. The DNA of foreign viruses and plasmids is incorporated into the CRISPR locus by the action of the Cas protein. Once incorporated into the CRISPR locus, a part of the foreign DNA functions as a template for producing crRNA. The produced crRNA forms a complex with tracrRNA to function as a guide strand, recruits the Cas protein to the foreign DNA complementary to the crRNA, and cleaves near the PAM sequence of the foreign DNA to eliminate the foreign DNA and exert an immune function. By focusing on the function of cleaving a predetermined sequence and designing a guide RNA that exhibits the functions of crRNA and tracrRNA to be complementary to the sequence of the target DNA, cleavage of the DNA can be induced at a desired position. Therefore, the advantage is the efficiency of obtaining a desired strain from a significantly smaller number of mutants compared to existing mutation introduction methods such as radiation, and it is used for various applications.

[0003] In typical genome editing technologies such as CRISPR / Cas9 and CRISPR / Cas12a, the sequence to be edited is primarily defined by a region called the spacer sequence of the gRNA or crRNA that constitutes it. The gRNA or crRNA spacer sequence is used as a target sequence, typically consisting of 20-24 bases that can be arbitrarily modified. In addition to this target sequence, for the nuclease Cas9 or Cas12a to function, a base sequence of approximately 2-4 bases called a PAM sequence must be present in the target nucleic acid sequence. Therefore, genome editing is performed by specifically recognizing a total base sequence of approximately 22-28 bases, combining the PAM sequence and the target sequence. While the target DNA sequence must contain a PAM sequence, the PAM sequence differs depending on the type of Cas protein family. By selecting the Cas protein that matches the target DNA sequence, genome editing of the desired sequence becomes possible.

[0004] The Cas protein, an RNA guide endonuclease, is hypothesized to originate from the IscB and TnpB proteins contained in transposons (Non-Patent Literature 1: J. Bacteriol. 198, 797-807 (2015)). The IscB and TnpB proteins, which are nucleases contained in transposons, have been reported to function as RNA guide endonucleases (Non-Patent Literature 2: Science 374, 57-65 (2021), Non-Patent Literature 3: Nature 599, 692-696 (2021)). Furthermore, it has been reported that complexes containing proteins classified as TnpB, which include a Ruv-C nuclease domain, and ωRNA molecules can be used for genome editing (Patent Document 1: International Publication No. 2022 / 159892). Research is actively being conducted on the functional analysis of TnpB and the identification of TnpB orthologues for use as new genome editing tools (Non-Patent Document 4: The CRISPR Journal. Jun 2023 232-242, Non-Patent Document 5: Nat Biotechnol., 2023), Non-Patent Document 6: Nature 620, 660-668, 2023, Non-Patent Document 7: Science Advances Sep 2023 vol. 9 Issue 39, Non-Patent Document 8: Nucleic Acids Research, gkad1053, 2023, Non-Patent Document 9: Proc Natl Acad Sci USA. 2023 Nov 28;120(48):e2308224120.). [Prior art documents] [Patent Documents]

[0005] [Patent Document 1] International Publication No. 2022 / 159892 [Non-patent literature]

[0006] [Non-Patent Document 1] J. Bacteriol. (2015) 198, 797-807 [Non-Patent Document 2] Science (2021) 374, 57-65 [Non-Patent Document 3] Nature (2021) 599, 692-696 [Non-Patent Document 4] The CRISPR Journal.(2023) vol. 6, No. 3, p.232-242 [Non-Patent Document 5] Nat Biotechnol., 2023, June 29, p.1-13 [Non-Patent Document 6] Nature, 2023, 620, 660-668 [Non-Patent Document 7] Science Advances Sep 2023 vol. 9 Issue 39, eadlk0171 [Non-Patent Document 8] Nucleic Acids Research, gkad1053,2023 [Non-Patent Document 9] Proc Natl Acad Sci USA. 2023 Nov 28;120(48):e2308224120. [Non-Patent Document 10] Int J Mol Sci., 2020 Apr 25;21(9):3038. doi: 10.3390 / ijms21093038. [Overview of the project] [Problems that the invention aims to solve]

[0007] The objective is to identify and utilize novel RNA guide nucleases applicable to genome editing technology. [Means for solving the problem]

[0008] The present inventors conducted intensive research with the aim of identifying novel proteins that do not form RNA guide endonuclease (hereinafter also referred to as RGN) complexes applicable to genome editing technology and are not annotated with TnpB. As a result, they discovered a TnpB-like RNA guide nuclease that is not classified as TnpB, leading to the present invention. Therefore, the present invention relates to the following: [1] (i) The following: YKGRTFNKMINNGSKGQYNxR-SxNxLKWRG (Sequence ID 109) [In the formula, x represents any amino acid] It contains a TnpB-like RNA guide nuclease with a length of 300-700 amino acids, (ii) A TnpB-like RNA guide nuclease complex comprising a guide RNA consisting of a guide sequence and a sequence containing a gRNA scaffold sequence. [2] The TnpB-like RNA guide nuclease complex described in item 1, wherein the TnpB-like RNA guide nuclease is a protein comprising an amino acid sequence selected from the group consisting of SEQ ID NOs: 16, 20, 22, 24, 26, 40, 50, 52, 54, and 56. [3] The TnpB-like RNA guide nuclease complex according to item 1, wherein the TnpB-like RNA guide nuclease is a mutant protein having an amino acid substitution mutation with respect to the protein described in item 2, and is a mutant protein having RNA guide nuclease activity. [4] The TnpB-like RNA guide nuclease complex described in item 3, wherein the mutant protein comprises an amino acid sequence having at least 75% sequence identity with the amino acid sequence of the TnpB-like RNA guide nuclease complex described in item 2. [5] The mutant protein is a protein that has structural homology with the original protein in the RuvC domain, where structural homology is defined as a TnpB-like RNA guide nuclease complex as described in item 3, where the root mean square deviation (RMSD) between the original protein backbone and the mutant backbone is 1.5 or less and the TM value is 0.9 or more when the structural alignment of the original protein and the mutant protein is performed using PyMOL software. [6] The TnpB-like RNA guide nuclease complex described in any one of items 1 to 5, wherein the TAM sequence of the TnpB-like RNA guide nuclease is a sequence containing TCAT. [7] The TnpB-like RNA guide nuclease complex described in item 6, wherein the TAM sequence includes TTCAT. [8] The TnpB-like RNA guide nuclease complex according to any one of items 1 to 7, wherein the TnpB-like RNA guide nuclease is a TnpB-like RNA guide nuclease consisting of the amino acid sequence of SEQ ID NO: 22, or a mutant protein thereof, and the gRNA scaffold sequence includes the sequence of SEQ ID NO: 58, or a sequence having at least 70% identity with said sequence. [9] The TnpB-like RNA guide nuclease complex described in item 8, wherein the mutant protein has an amino acid substitution at at least one position selected from the group consisting of Q at position 284, R at position 458, K at position 89, S at position 72 and / or S at position 75, R at position 291, S at position 201, L at position 278, and S at position 315.

[10] In the mutant protein, If there is an amino acid substitution at position 284, it is either Q284R or Q284K. If there is an amino acid substitution at position 89, it is K89R. If there is an amino acid substitution at position 72 or 75, the amino acids are S72H, S72N, S72Q, S72K, S72R, S75H, S75N, S75Q, S75K, or S75R. If there is an amino acid substitution at position 291, it is either R291E or R291K. If there is an amino acid substitution at position 201, it is either Q201H or Q201S. When there is an amino acid substitution at position 278, it is L278K or L278R, When there is an amino acid substitution at position 315, it is S315A or S135V, The TnpB-like RNA-guided nuclease complex according to item 9.

[11] The TnpB-like RNA-guided nuclease is a TnpB-like RNA-guided nuclease consisting of the amino acid sequence of SEQ ID NO: 22, or a mutant protein thereof, and the gRNA scaffold sequence is a nucleic acid consisting of a base sequence having 90% sequence identity to SEQ ID NO: 117 (gRNA1), SEQ ID NO: 118 (gRNA2), SEQ ID NO: 119 (gRNA3), or the said sequence. The TnpB-like RNA-guided nuclease complex according to any one of items 1 to 10.

[12] A method for cleaving a target DNA using the TnpB-like RNA-guided nuclease complex according to any one of items 1 to 11.

[13] The cleavage method is A nucleic acid containing an expression cassette encoding the TnpB-like RNA-guided nuclease, and Introducing into the cell a nucleic acid containing an expression cassette encoding the guide RNA. The cleavage method according to item 12.

[14] Using the cleavage method as a genome editing method including sequence-specific knockout and sequence-specific knock-in. The cleavage method according to item 12.

[15] (i) A TnpB-like RNA-guided nuclease having an inactivating mutation in the TnpB-like RNA-guided nuclease of the TnpB-like RNA-guided nuclease complex according to any one of items 1 to 11, and a transcription regulatory domain is fused, and (ii) A guide RNA consisting of a sequence containing a guide sequence and a gRNA scaffold sequence. An inactivating mutation-containing TnpB-like RNA-guided nuclease complex. /

[16] A method for transcription regulation of a target DNA using the inactivating mutation-containing TnpB-like RNA-guided nuclease complex according to item 15.

[17] The transcription regulation method is An expression cassette encoding the inactivating mutation-containing TnpB-like RNA-guided nuclease, and A method for regulating the transcription of target DNA as described in item 15, comprising introducing an expression cassette encoding the guide RNA into a cell.

[18] Genome editing kits containing a TnpB-like RNA guide nuclease complex as described in any one of items 1-11.

[19] A genome editing kit as described in item 18, comprising a guide RNA containing a gRNA scaffold sequence, and a guide sequence designed from a cleavage target sequence.

[20] The genome editing kit described in item 18, wherein the TAM sequence of the TnpB-like RNA guide nuclease is a sequence containing TCAT.

[21] The genome editing kit described in any one of items 18 to 20, which is used for sequence-specific knockout and sequence-specific knock-in. [Effects of the Invention]

[0009] According to the present invention, a novel RNA-guided endonuclease enables the cleavage of target nucleic acids, which can be applied to genome editing technology. [Brief explanation of the drawing]

[0010] [Figure 1] Figure 1 shows the phylogenetic relationships of T proteins into five groups (clades 1-5). The phylogenetic tree was constructed using the neighbor-joining method with the protein sequences listed in Table 1. [Figure 2] Figure 2 shows the 3' end sequences of genes encoding TnpB-like RNA guide nucleases belonging to each clade, based on sequence analysis results. Focusing on the 3' end boundaries of regions with high homology between each clade, these boundary regions are determined as TEM sequences. [Figure 3] Figure 3 shows the results of separating a TnpB-like RNA guide nuclease solution obtained by column purification after expressing recombinant His-MBP-TEV fused TnpB-like RNA guide nuclease in Escherichia coli, and staining it with Coomassie blue (A:T2, B:T3, C:T4, D:T5, E:T6, F:T7, G:T8-1, H:T8-2). [Figure 4] Figure 4 shows the results of next-generation sequencing analysis of RNA that was extracted from the TnpB-like RNA guide nuclease. It was shown that the RNA that was complexed with the TnpB-like RNA guide nuclease partially overlapped with the open reading frame of the TnpB-like RNA guide nuclease gene. [Figure 5] Figure 5 shows the results of analyzing cleaved plasmids after applying a TnpB-like RNA guide nuclease complex with guide RNA to a plasmid library containing random TAMs and target sequences. Plasmids containing TCAT(T3), TTAT(T5 and T7), and TTCAT(T6, T8-1, T8-2) were cleaved in the random sequence, and these sequences became the TAM sequences. [Figure 6] Figure 6 shows the results of in vitro investigation of cleavage activity when double-stranded DNA with different target sequences, amplified from purified plasmids containing TAM sequences, was treated with a complex of a guide RNA corresponding to the target sequence and a TnpB-like RNA guide nuclease. [Figure 7] Figure 7 shows the results of in vitro cleavage activity analysis of purified plasmids containing different TAM sequences and different target sequences, by treating them with a complex of a target-specific guide RNA and TnpB-like RNA guide nuclease (T5 and T8-2). The types of RNPs treated and the TAM sequences used are shown at the top. No RNP indicates a negative control. [Figure 8] Figure 8 shows the results of in vitro investigations into the cleavage activity of purified plasmids containing TAM sequences and target sequences for T5 and T8, respectively, by treating them with complexes of guide RNA and TnpB-like RNA guide nucleases (T5 and T8-2) according to the target sequence, under various temperature changes. [Figure 9]Figure 9 is a graph showing the genome editing efficiency when HEK293FT is transfected with an expression vector encoding T8-1 protein or T8-2, and an expression vector encoding guide RNA. The vertical axis represents the absolute value of genome editing efficiency, and the horizontal axis represents the conditions for transformation with T8-1 gRNA1-5 and T8-2 gRNA1-5, respectively. The Control column shows the genome editing efficiency when T8 expression and guide DNA expression plasmids are not included. [Figure 10] Figure 10 shows the nucleic acid sequences of the genome near the cleavage site (5' to the left, 3' to the right) when HEK293FT is transfected with an expression vector encoding the T8-1 protein and an expression vector encoding a guide RNA. The sgRNA indicates the target sequence and recognizes the complementary strand of the indicated DNA. The TAM sequence is located on the 3' side of the target sequence and is not included in this figure. The dashed line indicates the expected location of the double-strand break. [Figure 11] Figure 11 shows the editing efficiency when modified T8-2 is introduced using gRNA original, which has a full-length T8-2 guide RNA scaffold sequence of 152 nt as the guide RNA, or ddHDV, which has only the HDV sequence removed, gRNA1 in which the T8-2 guide RNA scaffold sequence has been shortened to 131 nt, gRNA2 in which the guide RNA scaffold sequence has been shortened to 111 nt, gRNA3 in which it has been shortened to 100 nt, and gRNA4 in which it has been shortened to 77 nt. [Figure 12] Figure 12 shows the effect of point mutations in the modified T8-2. [Figure 13] Figure 13 shows that the production of the target protein is increased by an inactivating mutant-containing TnpB-like RNA guide nuclease complex, which includes a TnpB-like RNA guide nuclease in which the P300 core domain is fused to an inactivating mutant-containing T8-2, and a guide RNA. [Modes for carrying out the invention]

[0011] One aspect of the present invention relates to a TnpB-like RNA guide nuclease complex, a kit for expressing the complex, a method for cleaving target DNA using such a TnpB-like RNA guide nuclease complex, and a genome editing method.

[0012] [TnpB-like RNA guide nuclease complex] The TnpB-like RNA guide nuclease complex according to the present invention is as follows: (1) TnpB-like RNA guide nuclease, and (2) gRNA consisting of a sequence including a guide sequence and a gRNA scaffold sequence The complex may cleave a predetermined DNA in a test tube, or it may cleave a predetermined position of DNA complementary to the guide RNA sequence by forming the complex in a cell. The complex can be formed by introducing a vector containing (1) a TnpB-like RNA guide nuclease and (2) an expression cassette encoding the guide RNA, or by introducing messenger RNA or the guide RNA itself into a cell and expressing it, or a complex that has been formed in vitro may be introduced into the cell using techniques such as liposomes or electroporation.

[0013] [TnpB-like RNA guide nuclease] TnpB-like RNA guide nucleases (hereinafter also called T proteins) that form the TnpB-like RNA guide nuclease complex contain one or more of the following characteristics: (i) The following regular expression array: YKGRTF-[NS]-[KR]-[LM]-x-[AN]-xG-[AS]-[KR]-[GS]-QYx(2)-R-[AS]-x-[DN]-xLxWxG (Sequence ID 1) (In the formula, [NS] represents N or S, [KR] represents either K or R. [LM] represents L or S, x represents any amino acid, [AN] represents A or N, [AS] represents A or S, [KR] represents either K or R. [GS] represents G or S, x(2) represents two arbitrary amino acids, [AS] represents A or S, [DN] represents either D or N. Includes; (ii) comprising at least one domain selected from the group consisting of RuvCI, bridge helix, RuvCII, RuvCIII, and Wedge; (iii) Having a length of 300 to 700 residues, preferably 500 to 675 residues, more preferably 550 to 650; and (iv) Possesses RNA guide nuclease activity.

[0014] The regular expression sequence in this invention is a conserved sequence in TnpB-like RNA guide nucleases of clades 1 to 4. The regular expression sequence is a 29-amino acid sequence contained in the RuvC domain (more specifically, RuvCIII) of the TnpB-like RNA guide nuclease. By performing a database search on this regular expression sequence, all T proteins can be specifically identified.

[0015] More specifically, the TnpB-like RNA guide nuclease has the following regular expression sequence: (i) The following: YKGRTFNKMINNGSKGQYNxR-SxNxLKWRG (Sequence ID 109) (In the formula, x independently represents any amino acid.) This relates to TnpB-like RNA guide nucleases that contain and have a length of 300-700 amino acids. Such regular expression sequences refer to TnpB-like RNA guide nucleases included in clade 1 that do not contain T11, T16, or T18. Specifically, this may refer to TnpB-like RNA guide nucleases represented, for example, T6, T8-1 to T8-4, T15, T17, and T21-1 to T21-4.

[0016] The bridge helix, RuvCI, RuvCII, and RuvCIII are thought to be involved in the endonuclease domain, and Cas, IscB, and TnpB also have similar domains. On the other hand, the regular expression sequence in RuvCIII is shared only with TnpB-like RNA guide nucleases and does not match Cas9, Cas12, IscB, and TnpB. The TnpB-like RNA guide nuclease according to the present invention may include all domains of the bridge helix, RuvCI, RuvCII, and RuvCIII from the viewpoint of exhibiting RNA guide nuclease activity. Furthermore, it may be identified by including common sequences in each domain.

[0017] The RNA guide nuclease activity of TnpB-like RNA guide nucleases recognizes a 3-7 base sequence, particularly a 4- or 5-base sequence, called a TAM (Transposon Associated Motif) sequence, and cleaves the target DNA sequence downstream of it. Therefore, the target sequence to which the guide RNA binds is designed to be adjacent to the TAM sequence, and the TnpB-like RNA guide nuclease cleaves the target nucleic acid containing both the target sequence and the TAM sequence. While the TAM sequence may differ for each TnpB-like RNA guide nuclease, the same TAM sequence can usually be used within the same clade (Figure 5). Furthermore, the TAM sequence can be determined by a method well known in this art: a random TAM library in which a random region of several bases, for example 7 bases, is linked to the target sequence is treated with a complex of the target sequence guide RNA and a TnpB-like RNA guide endonuclease, and the cleavage is confirmed to determine the TAM sequence. As an example, the typical TAM sequences recognized by the TnpB-like RNA guide nucleases of each clade are as follows: [Table 1]

[0018] TnpB-like RNA guide nucleases can also be specifically referred to as T proteins (e.g., T1-T21) and can be classified phylogenetically into clades 1 to 5. Phylogenetic classification can be performed using methods well known in this field. For example, homologs of TnpB-like RNA guide nucleases can be searched using NCBI's PHI-BLAST by using a regular expression of the amino acid sequence (sequence number 1) or a regular expression of clade 1 (sequence number 109) as a seed query. The results can then be further searched for and collected proteins with sequence homology using BLASTP software, and alignment using ClustalW or similar software can be performed to create a phylogenetic tree from the alignment scores. Alternatively, the amino acid sequence of the T8-2 protein (sequence number 22) can be used as a seed query to detect it using NCBI BLASTP for non-redundant protein sequences, and the classification can be determined based on phylogenetic tree analysis of protein sequences using the neighbor-joining method. As another example, using the three-dimensional structure of a portion of the RuvC III domain as a query in FoldSeek, the protein family can also be classified based on the level of its RMSD score. This invention relates particularly to T proteins of clades 1 to 4. The proteins included in clade 1 are T6, T8-1 to T8-4, T11, T15, T16, T17, T18, and T21-1 to T21-4. In particular, this concerns TnpB-like RNA guide nucleases represented by T6, T8-1 to T8-4, T15, T17, and T21-1 to T21-4, characterized by the absence of T11, T16, and T18. The proteins included in clade 2 are T2, T3, and T19. The proteins included in clade 3 are T4-1, T4-2, and T5. The proteins included in clade 4 are T1, T13, and T14. The proteins included in clade 5 are T7, T9, T10, T12, and T20. Phylogenetically, clade 5 is most closely related to known TnpB, followed by clade 4, clade 3, clade 2, and then clade 1 in increasing order of relative distance.

[0019] The amino acid sequences of proteins belonging to clade 1 and the nucleotide sequences of ORFs are as follows: [Table 2]

[0020] The amino acid sequences of proteins belonging to clade 2 and the nucleotide sequences of ORFs are as follows: [Table 3]

[0021] The amino acid sequences of proteins belonging to clade 3 and the nucleotide sequences of ORFs are as follows: [Table 4]

[0022] The amino acid sequences of proteins belonging to clade 4 and the nucleotide sequences of ORFs are as follows: [Table 5]

[0023] The amino acid sequences of proteins belonging to clade 5 and the nucleotide sequences of ORFs are as follows: [Table 6]

[0024] In one aspect of the present invention, the TnpB-like RNA guide nuclease of the present invention is a protein having an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 49, 50, 52, 54, and 56, or Regarding that mutated protein. A mutant protein is a protein that has an amino acid sequence with at least 50% sequence homology or identity with the original protein's amino acid sequence and possesses TnpB-like RNA guide nuclease activity. Such mutant proteins have the following canonical sequence in the RuvC domain: YKGRTF-[NS]-[KR]-[LM]-x-[AN]-xG-[AS]-[KR]-[GS]-QYx(2)-R-[AS]-x-[DN]-xLxWxG (Sequence ID 1) (In the formula, [NS] represents N or S, [KR] represents either K or R. [LM] represents L or S, x represents any amino acid, [AN] represents A or N, [AS] represents A or S, [KR] represents either K or R. [GS] represents G or S, x(2) represents two arbitrary amino acids, [AS] represents A or S, [DN] represents either D or N. It is characterized by containing a specific component, and the size of the RuvC domain is specified to be 200 to 350 in total length, preferably 250 to 300. Furthermore, such mutant proteins have steric structural homology, and more preferably structural homology in the RuvC domain.

[0025] The mutant protein may be identified with 40% or more, 45% or more, 50% or more, 55% or more, 60% or more, 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 99% or more homology and / or identity with respect to the amino acid sequence of the original protein. In particular, it is especially preferable that the mutant protein belongs to the same clade as the original protein. If they belong to the same clade, the homology or identity is usually preferably 50% or more, 60% or more, 70% or more, or 80% or more.

[0026] Examples of TnpB-like RNA guide nucleases belonging to clade 1 include the following: (1) A protein consisting of an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 16, 20, 22, 24, 26, 32, 40, 42, 46, 50, 52, 54, and 56, in particular a protein consisting of an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 16, 20, 22, 24, 26, 40, 50, 52, 54, and 56, or (2) An amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 16, 20, 22, 24, 26, 32, 40, 42, 46, 50, 52, 54, and 56, in particular an amino acid sequence having at least 50%, 60%, 70%, 75%, 80%, or 90% sequence homology or identity with the amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 16, 20, 22, 24, 26, 40, 50, 52, 54, and 56, and the sequence is a protein having TnpB-like RNA guide nuclease activity. Regarding this, for sequence identity, T11, T16, and T18 can be appropriately selected so as not to be included.

[0027] Examples of TnpB-like RNA guide nucleases belonging to clade 2 include the following: (1) A protein consisting of an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 4, 6, 8, and 48, or (2) A protein having TnpB-like RNA guide nuclease activity, comprising an amino acid sequence having at least 40%, 50%, or 60% sequence homology or identity with an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 4, 6, 8, and 48, wherein the sequence is a protein having TnpB-like RNA guide nuclease activity. Regarding.

[0028] Examples of TnpB-like RNA guide nucleases belonging to clade 3 include the following: (1) A protein consisting of an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 10, 12, and 14, or (2) A protein having TnpB-like RNA guide nuclease activity, comprising an amino acid sequence having at least 80% sequence homology or identity with an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 10, 12, and 14. Regarding.

[0029] Examples of TnpB-like RNA guide nucleases belonging to clade 4 include the following: (1) A protein consisting of an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 2, 36, and 38, or (2) A protein having TnpB-like RNA guide nuclease activity, comprising an amino acid sequence having at least 60% sequence homology or identity with an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 2, 36, and 38. Regarding.

[0030] In one embodiment, the TnpB-like RNA guide nuclease of the present invention is as follows: (1) A protein encoded by a nucleic acid sequence selected from the group consisting of the nucleic acid sequences of SEQ ID NOs: 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 51, 53, 55, and 57, or (2) A protein comprising a nucleic acid sequence that has at least 60% sequence identity with a nucleic acid sequence selected from the group consisting of the nucleic acid sequences of SEQ ID NOs: 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 51, 53, 55, and 57, wherein the protein has TnpB-like RNA guide nuclease activity. Regarding.

[0031] In this specification, "homology" of two amino acid sequences refers to the ratio of identical or similar amino acid residues that appear at corresponding locations when the two amino acid sequences are aligned, and "identity" of two amino acid sequences refers to the ratio of identical amino acid residues that appear at corresponding locations when the two amino acid sequences are aligned. The "homology" and "identity" of two amino acid sequences can be determined, for example, using a program such as BLAST (Basic Local Alignment Search Tool) (Altschul et al., J. Mol. Biol., (1990), 215(3):403-10).

[0032] A mutant protein is a protein that has structural homology to the original protein in the RuvC domain, and may also have structural homology to the original protein in terms of its entire length. The position of the RuvC domain varies depending on the type of T protein, but for example, in the case of the T8-2 protein, R302 to E580 corresponds to the RuvC domain. For other T proteins, the RuvC domain corresponding to these positions can be determined. Structural homology can be determined by the RMSD value and / or TM value when aligned with the three-dimensional structure of the original protein. The RMSD value refers to the root mean square deviation of atomic positions. For example, if the root mean square deviation (RMSD) between the backbone of the original protein and the backbone of the mutant is 2.5A or less, preferably 1.5A or less, and more preferably 1A or less, it is highly likely that the mutant protein will produce an effect equivalent to that of the original protein. The RMSD value can be determined using software known in the art, such as PyMOL software align. In addition to or instead of the RMSD value, the TM value (template modeling score) can also be used. Proteins with a TM value of 0.9 or higher, preferably 0.91 or higher, and more preferably 0.93 or higher, can be said to have structural homology. Therefore, variants can be identified by the RMSD value in addition to or instead of by sequence identity or homology. Specifically, the T8-2 protein belonging to clade 1 and the TnpB-like RNA guide nucleases belonging to clades 1-4 have an RMSD value of 1.9 or less and a TM value of 0.9 or more when structurally aligned using PyMOL software. In another example, when belonging to the same clade, for example, the T8-2 protein belonging to clade 1 and the TnpB-like RNA guide nuclease belonging to clade 1 have an RMSD value of 1.5 or less and a TM value of 0.91 or more when structurally aligned using PyMOL software. Since all proteins within a clade have TnpB-like RNA guide nuclease activity, if the RMSD between the original protein and the mutant protein is 1.9 or less, preferably 1.5 or less, more preferably 1 or less, and the TM value is 0.9 or more, preferably 0.91 or more, more preferably 0.93 or more, then the mutant protein is presumed to have equivalent activity to the original and can be used with genome editing tools in the procedure described herein.

[0033] In another example, the three-dimensional structure and homology of the T8-2 protein belonging to clade 1 and a known highly active TnpBRNA guide nuclease are compared. DraTnpB-AI (SEQ ID NO: 174) is an example of a known highly active TnpBRNA guide nuclease. When the three-dimensional structures are aligned, the amino acid positions that contribute to activity can be identified. In the amino acid sequence of the T8-2 protein (SEQ ID NO: 22), the amino acid positions corresponding to Q at position 284, R at position 458, K at position 89, S at position 72 and / or S at position 75, R at position 291, S at position 201, L at position 278, and S at position 315 may contribute to the activity of the TnpB-like RNA guide nuclease belonging to clade 1. The corresponding amino acid positions can be determined by structural comparison or amino acid sequence comparison. In one embodiment, the present invention relates to an amino acid sequence of a TnpB-like RNA guide nuclease belonging to clade 1, wherein the amino acid sequence of sequence number 22 includes a sequence having a substitution at a corresponding position in at least one selected from the group consisting of Q at position 284, R at position 458, K at position 89, S at position 72 and / or S at position 75, R at position 291, S at position 201, L at position 278, and S at position 315. These mutations are particularly preferably combinations of two positions or combinations of three positions. As an example, mutations at the positions of K at position 89 and Q at position 284, K at position 89 and R at position 291, Q at position 284 and R at position 291, or combinations of K at position 89, Q at position 284 and R at position 291 are preferred.

[0034] Such substitutions include Q284R or Q284K if the substitution is at position 284 or a corresponding position, K89R if the substitution is at position 89 or a corresponding position, S72H, S72N, S72Q, S72K, S72R, S75H, S75N, S75Q, S75K, or S75R if the substitution is at position 72 or 75 or a corresponding position, R291E or R291K if the substitution is at position 291 or a corresponding position, Q201H or Q201S if the substitution is at position 201 or a corresponding position, L278K or L278R if the substitution is at position 278 or a corresponding position, and S315A or S135V if the substitution is at position 315 or a corresponding position. Having these substitutions can enhance TnpB-like RNA guide nuclease activity. More specifically, from the viewpoint of achieving high genome editing efficiency, substitution of K89R, Q284K, Q284R, and R291K, or combinations thereof, is preferred. For example, K89R and Q284K, K89R and R291K, Q284K and R291K, or K89R, Q284K, and R291K are preferred (Figure 12).

[0035] The TnpB-like RNA guide nuclease or its mutant protein of the present invention exhibits TnpB-like RNA guide nuclease activity over a wide temperature range, for example, 20-50°C, but its activity is particularly high at 30-45°C, and more preferably around 37°C (Figure 8). Typical TnpB proteins and TnpB-derived proteins often exhibit maximum activity around 40-50°C, making it possible to use them efficiently at lower temperatures. Because it exhibits high activity not only around 37°C, the body temperature of homeothermic animals, but also at 25°C, it is particularly suitable for genome editing in living organisms such as plants, algae, mollusks, and microorganisms.

[0036] [Guide RNA (gRNA)] In this invention, the guide RNA is designed based on the DNA sequence of the cleavage target and the TnpB-like RNA guide nuclease to be used, and is configured to include a guide sequence and a gRNA scaffold sequence. The DNA of the cleavage target is selected under the condition that it includes the TAM sequence, since it is cleaved near the TAM sequence. A 14nt-25nt sequence adjacent to the 3' end of the TAM sequence can be selected as the guide sequence. The guide sequence may match the DNA sequence of the cleavage target, or it may contain one or more mismatched sequences. Even if the guide sequence does not perfectly match the target sequence, cleavage may still occur due to off-target effects. The gRNA scaffold must use a sequence that is available to the TnpB-like RNA guide nuclease being used. Typically, the 3' end boundary of the locus (i.e., the boundary between the gRNA scaffold and the guide sequence) can be determined by sequence comparison at the locus of each TnpB-like RNA guide nuclease. The last four bases of the 3' end of the gRNA scaffold sequence can be defined as the transposon end motif (TEM). The gRNA scaffold sequence can be 70–200 nt. The guide sequence may be designed to be adjacent to the 3' end of the scaffold sequence, i.e., the TEM sequence, or it may be an insertion of any sequence of 20–100 nt.

[0037] Within each clade, the gRNA scaffold sequences of TnpB-like RNA guide nucleases exhibit high homology / identity. For example, within the same clade, gRNA scaffold sequences can be used if they have 70% or more, preferably 80% or more, and more preferably 90% or more identity / homology. Representative gRNA scaffold sequences for each grade are as follows: [Table 7] The underlined parts represent the TEM sequence. Therefore, as the gRNA scaffold, any sequence of sequence numbers 58 to 62, or a sequence having 70% or more, preferably 80% or more, and more preferably 90% or more sequence identity with said sequence, can be used.

[0038] Based on sequence identity within a clade, gRNA scaffold sequences can identify regions that are more important for activity, and shortened sequences containing such important regions can be created. For example, shortened sequences of the T8-2 gRNA scaffold sequence include SEQ ID NOs. 117 (gRNA1), 118 (gRNA2), and 119 (gRNA3). Alternatively, shortened mutant sequences that have at least 90%, more preferably at least 92%, at least 93%, 95%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NOs. 119 (gRNA3) and are capable of functioning as an RNA scaffold can also be selected.

[0039] The gRNA, designed based on the TnpB-like RNA guide nuclease and the nucleic acid sequence of the cleavage target, can be expressed using methods well known in the art. For example, a nucleic acid containing an expression cassette in which DNA containing the gRNA sequence is positioned under the control of a promoter can be used. Such an expression cassette may contain a transcription termination sequence in its 3' region. The nucleic acid containing the expression cassette is loaded onto a vector or plasmid and introduced in vitro or into a designated cell to generate the gRNA.

[0040] [TnpB-like RNA guide nuclease containing inactivating mutations] TnpB-like RNA guide nucleases can be inactivated by substituting amino acid residues in their active domain. Examples of active domains include D308, E489, and D575 in T8-2; substituting residues at these positions inactivates the TnpB-like RNA guide nuclease. Inactivated TnpB-like RNA guide nucleases are recruited to target sequences by the action of guide RNA, but they do not exhibit RNA guide nuclease activity. Further fusion of a functional domain with an activating mutant TnpB-like RNA guide nuclease can regulate its function in the target sequence. The fused functional domain can be any domain, but examples include transcriptional regulatory domains such as p300, DMNT3, KRAB, SRDX, VPR, and VP64. In another example, any functional domain such as APOBEC, AID, TadA, Dda, FokI, M-MLV-RTase, or GFP can be used, enabling not only epigenetic states but also base modification, reverse transcription, phosphate bond cleavage, and photomanipulation such as fluorescent proteins. RNA guide nucleases containing inactivating mutations fused with functional domains are also widely known in the CRISPR-Cas system (Non-patent Literature 10: Int J Mol Sci., 2020 Apr 25;21(9):3038. doi: 10.3390 / ijms21093038), and similar functional regulation is possible with TnpB-like RNA guide nucleases.

[0041] An inactivating mutant-containing TnpB-like RNA guide nuclease fused with a functional domain forms a complex with a guide RNA consisting of a guide sequence and a gRNA scaffold sequence, and exerts a function corresponding to the functional domain. When a transcription-promoting domain, such as all or part of the functional domains of p300, VPR, or VP64, is fused as the functional domain, transcription is promoted in the region recruited by the guide RNA. On the other hand, when a transcription-repressing domain, such as all or part of the functional domains of DMNT3, KRAB, or SRDX, is fused as the functional domain, transcription is repressed in the region recruited by the guide RNA. Therefore, the present invention may also relate to a method for regulating the transcription of target DNA using an inactivating mutant-containing TnpB-like RNA guide nuclease fused with such a functional domain.

[0042] [Target-specific cleavage method] Another aspect of the present invention relates to a method for cleaving target DNA using a TnpB-like RNA guide nuclease complex according to the present invention. More specifically, the following: (i) The TnpB-like RNA guide nuclease according to the present invention, (ii) Guide RNA consisting of a sequence including a guide sequence and a gRNA scaffold sequence This invention relates to a method for cleaving target DNA using a complex containing the TnpB-like RNA guide nuclease. The complex may be formed in vitro and introduced into cells using electroporation or liposomes to cleave the target DNA, or the target DNA may be cleaved in vitro. Alternatively, the complex may be mounted on an expression cassette, and the gene may be introduced into cells using a vector, messenger RNA, or guide RNA itself. By expressing the TnpB-like RNA guide nuclease and the guide RNA in the cell, the DNA of the genome may be cleaved. Such cleavage methods can be performed using a kit for target nucleic acid cleavage that includes an expression cassette.

[0043] Gene transfer can be performed using methods known in this field. The transfer method is not particularly limited and can be appropriately selected depending on the type of substance to be transferred and the target of the transfer. Transfer methods can be broadly classified into direct methods and methods using viral vectors. Direct methods include electroporation, liposome methods, particle gun methods using gold particles, and whisker methods. Methods using viral vectors can use adenoviruses, adeno-associated viruses, lentiviruses, Agrobacterium, tobacco mosaic virus (TMV), etc., depending on the host species.

[0044] Target-specific cleavage within cells can be repaired by genomic DNA repair mechanisms. This can lead to non-homologous end joining or homologous recombination, enabling gene knockout or knock-in. Therefore, the target-specific cleavage method in this invention can also be called a genome editing method. The genome editing method in this invention may be performed on cultured cells, cultured tissues, or cells of living organisms. The target species is not particularly limited, but genome editing is possible for any cells of bacteria, archaea, and eukaryotes. Eukaryotes may include any cells such as plant cells, insect cells, and animal cells; for example, genome editing is possible for any mammalian cell, and genome editing is possible for both human and non-human animal cells. The T protein according to this invention is characterized by a low optimal temperature. Therefore, genome editing can be performed at temperatures of 10-30°C.

[0045] [Target-Specific Protein Recruitment] Another aspect of the present invention relates to a method for recruiting a specific protein to target DNA using a TnpB-like RNA guide nuclease complex according to the present invention, which has been modified so as not to exhibit TnpB-like RNA guide nuclease activity. As a complex modified to not exhibit TnpB-like RNA guide nuclease activity, the activity of the TnpB-like RNA guide nuclease may be inactivated, or the length of the guide RNA may be adjusted so that the TnpB-like protein does not exhibit DNA cleavage activity, while the guide RNA containing a sequence of a length that specifically binds to the target is used. As such a guide RNA, a guide sequence of a target sequence of about 10 bases may be used. More specifically, the TnpB-like RNA guide nuclease complex used in a method for recruiting a specific protein to target DNA is as follows: (i) A TnpB-like RNA guide nuclease according to the present invention, in which the amino acid residue of the active site is replaced with another amino acid, (ii) Guide RNA consisting of a sequence including a guide sequence and a gRNA scaffold sequence (iii) Specific proteins that bind to or interact with TnpB-like RNA guide nucleases This includes the following. By expressing or introducing such a complex into cells, the complex is recruited near the site of the guide sequence without cleavage, and specific proteins bound to the complex can be recruited. The active site of the mutant TnpB-like RNA guide nuclease described in (i) is a typical DED active site contained in the RuvC domain. In the case of T8-2, the activity of the TnpB-like guide nuclease can be inactivated by introducing mutations at D308, E489, and D575. Mutations can be appropriately selected within the range that can inactivate the activity, but one example is a mutation to alanine. (ii) The RNA scaffold sequence of the guide RNA described can have binding sites for other proteins, such as the MS2 sequence, inserted. (iii) The specific proteins described are proteins that modify the DNA or epigenetic state; specifically, all proteins commonly referred to as epigenetic factors are available. More specifically, these include p300, DMNT3, KRAB, SRDX, VPR, VP64, etc. In addition to the epigenetic state, these also modify bases, reverse transcription, phosphate bond cleavage, and the binding of photomanipulation tools such as fluorescent proteins (specifically APOBEC, AID, TadA, Dda, FokI, M-MLV-RTase, GFP, etc.), but are not particularly limited to these. By recruiting such specific proteins to a target DNA region, the epigenetic state of the desired DNA region can be altered.

[0046] [kit] Another aspect of the present invention may relate to a kit for target nucleic acid cleavage in cells or a kit for genome editing. (i) An expression cassette of a TnpB-like RNA guide nuclease according to the present invention, (ii) An expression cassette of guide RNA having a sequence containing a gRNA scaffold sequence This includes the following. The guide RNA expression cassette is provided with an insertable guide sequence. The guide RNA expression cassette can be prepared by designing a cleavage target sequence and introducing the guide sequence. The TnpB-like RNA guide nuclease expression cassette and the guide RNA expression cassette with the introduced guide sequence can be incorporated into a plasmid or vector, for example, and used for gene delivery into cells.

[0047] An expression cassette can typically be created by including a promoter sequence and placing a sequence encoding a TnpB-like RNA guide nuclease or guide RNA under its control. The promoter can transiently or constitutively control the expression of downstream sequences. The expression cassette may be contained on a single polynucleotide or on different polynucleotides. The expression cassette may contain elements that contribute to expression, such as a terminator, as well as elements necessary for plasmid or vector preparation, such as signal sequences, tag sequences, reporter sequences, multi-cloning sites, drug resistance genes, and origins of replication, as long as the activity of the TnpB-like RNA guide nuclease is not impaired. From the viewpoint of ensuring that the expressed protein acts in the nucleus, it is desirable that a nuclear localization signal be attached to the signal sequence. One or more nuclear localization signals may be arranged in series. The nuclear localization signal can be appropriately selected depending on the species. As a result, the TnpB-like RNA guide nuclease expressed in the cell can translocate into the nucleus and, in cooperation with the guide RNA, can cleave the target nucleic acid sequence.

[0048] The genome editing kit according to the present invention may further include, as necessary, other materials, reagents, and equipment required to carry out the genome editing method of the present invention, such as nucleic acid introduction reagents and buffers. Other materials required to carry out the genome editing method of the present invention include a TnpB-like RNA guide nuclease expression cassette and / or a guide RNA expression cassette, as well as a donor polynucleotide. By introducing the donor polynucleotide into the nucleus, it becomes possible to knock in the donor polynucleotide at the cleavage site created by the CRISPR / Cas system. The donor polynucleotide causes homologous recombination at the cleavage site by positioning homologous sequences to the 5' and 3' sequences of the cleavage site at the 5' and 3' ends of the introduced sequence, respectively.

[0049] All references made herein are incorporated herein by citation in their entirety.

[0050] The present invention will be described in more detail below with reference to examples, but these examples are merely illustrative examples for explanatory purposes, and the present invention is not limited in any sense to these examples. [Examples]

[0051] Example 1: Identification of TnpB-like RNA guide nuclease homologs Homologs of TnpB-like RNA guide nucleases were detected using NCBI BLASTP and TBLASTN for non-redundant protein sequences in the NCBI database, with the amino acid sequence of the T8-2 protein (SEQ ID NO: 22) used as a seed query. Similar searches were also performed in the MGNIFY database in addition to NCBI. The nucleotide sequences of the loci corresponding to each obtained TnpB-like RNA guide nuclease were obtained from the NCBI Genbank database. Based on phylogenetic analysis of protein sequences using the neighbor-joining method, TnpB-like RNA guide nucleases were classified into five groups (clades 1-5) (Figure 1). The right-hand boundary of the locus (i.e., the boundary between the guide RNA (gRNA) scaffold and the guide sequence) was determined by aligning the genomic sequences of the loci encoding the T proteins listed above using ClustalW (https: / / www.genome.jp / tools-bin / clustalw) and comparing their 3' ends. From the alignment, the guide-scaffold boundary was identified as the downstream location where sequence conservation rapidly declines (Figure 2). A sequence motif consisting of four nucleotides located at the 3' end of the scaffold was defined as the transposon-encoded motif (TEM).

[0052] By aligning the amino acid sequences of TnpB-like RNA guide nucleases contained in clades 1-4, we determined the conserved sequences (regular expressions) within these sequences: YKGRTF-[NS]-[KR]-[LM]-x-[AN]-xG-[AS]-[KR]-[GS]-QYx(2)-R-[AS]-x-[DN]-xLxWxG (Sequence ID 1) (In the formula, [NS] represents N or S, [KR] represents either K or R. [LM] represents L or S, x represents any amino acid, [AN] represents A or N, [AS] represents A or S, [KR] represents either K or R. [GS] represents G or S, x(2) represents two arbitrary amino acids, [AS] represents A or S, [DN] represents either D or N. By performing a database search using PHI-BLAST or similar tools on this regular expression sequence, it is possible to specifically identify all clades 1-4 of the T protein.

[0053] Example 2: Construction of a TnpB-like RNA-guided nuclease expression vector To produce recombinant TnpB-like RNA guide nuclease as a maltose-binding protein-TEV protease cleavage site (His-MBP-TEV) fusion protein in E. coli, a synthetic DNA fragment containing the His-MBP sequence was cloned between the NcoI and NdeI cleavage sites of pET28b using the NEBuilder HiFi DNA Assembly Kit (New England Biolabs). This was designated as the pAN36 vector. The resulting plasmid was named pAN36. All genes encoding the TnpB-like RNA guide nuclease and their associated gRNA scaffolds were synthesized by Integrated DNA Technology (IDT) (Table 8). For some synthetic TnpB-like RNA guide nuclease genes, the 5' region of the coding sequence that did not overlap with the gRNA scaffold was codon-optimized for protein expression in E. coli. These DNA fragments were cloned into the pAN36 vector under the T7 promoter. In the resulting vector, the sequence encoding the TnpB-like RNA guide nuclease was fused in-frame to the N-terminal His-MBP-TEV sequence. A description of the His-MBP-TEV-TnpB-like RNA guide nuclease expression vector is shown in Table 8. [Table 8]

[0054] Example 3: Expression and purification of TnpB-like RNA guide nuclease-RNA complex in Escherichia coli. To express recombinant His-MBP-TEV fusion TnpB-like RNA guide nuclease, Escherichia coli Rosetta2(DE3)pLysS strain (Novagen) was transformed with an expression vector for His-MBP-TEV fusion TnpB-like RNA guide nuclease. The Escherichia coli cells were cultured overnight at 37°C in Luria-Bertani (LB) medium supplemented with 50 μg / ml kanamycin and 34 μg / ml chloramphenicol. 10 ml of the overnight culture was inoculated into 1 liter of LB medium supplemented with 50 μg / ml kanamycin and 34 μg / ml chloramphenicol, and the optical density (OD) was measured.600 The cells were grown until the ratio reached 0.6. Then, 0.25 mM IPTG was added to induce gene expression, and the cells were grown at 18°C ​​for 20 hours. The cells were harvested by centrifugation and stored at -70°C until use.

[0055] All subsequent purification steps were performed at 4°C. The cells were resuspended in buffer A (50 mM Tris-HCl, pH 8.0, 500 mM NaCl, 5% (v / v) glycerol, and 25 mM imidazole) supplemented with bovine DNAse I (Fujifilm Wako Pure Chemical Industries), chicken lysozyme (Fujifilm Wako Pure Chemical Industries), and protease inhibitors (phenylmethylsulfonyl fluoride and Roche cOmplete ethylenediaminetetraacetic acid-free). After incubation for 30 minutes, the cells were disrupted on ice by sonication (ULTRASONIC DISRUPTOR UD-211, TOMY). After removing cell debris by centrifugation at 40,000 g for 30 minutes, the supernatant was filtered through a 0.45 μm polyvinylidene fluoride (PVDF) membrane and bound in batches to 1 ml of Ni-Sepharose 6 Fast Flow Resin (Cytiva), which had been equilibrated with buffer A for 1 hour. This resin was packed into an Econo-Pac chromatography column (Bio-Rad) and washed first with 14 ml of buffer B (50 mM Tris-HCl, pH 8.0, 1 M NaCl, 5% (v / v) glycerol, 25 mM imidazole), followed by 10 ml of buffer A. The bound proteins were eluted with 3.5 ml of buffer C (50 mM Tris-HCl, pH 8.0, 1 M NaCl, 5% (v / v) glycerol, 300 mM imidazole). Proteins in the eluate were separated by 5-20% (w / v) polyacrylamide gel SDS-PAGE, and the gel was stained with Coomassie brilliant blue (Figure 3: SDS-PAGE gel). The peak fraction containing the fusion protein was transferred to a nuclease-free tube, rapidly frozen with liquid nitrogen, and stored at -70°C until use. The resulting TnpB-like RNA guide nuclease ribonucleoprotein (hereinafter referred to as RNP) samples were used for nucleic acid extraction and dsDNA cleavage analysis. For in vitro double-strand DNS cleavage analysis, TnpB-like RNA guide nuclease (RNP) samples were further purified using a HiLoad 16 / 600 Superdex 200 pg column in Buffer D [25 mM Tris-HCl pH 8.0, 150 mM NaCl, 5 mM MgCl2, 1% (v / v) glycerol, 1 mM DTT]. The pooled fractions were concentrated by ultrafiltration and stored at -80°C until use.

[0056] Example 4: Extraction and analysis of gRNA bound to TnpB-like RNA guide nuclease To extract RNA bound to the TnpB-like RNA guide nuclease RNP, 180 μl of the peak fraction containing RNP was vigorously mixed with 540 μl of TRI Regent (Molecular Research Center, Inc.) and 108 μl of chloroform for 15 seconds. After incubation at room temperature for 5 minutes, the mixture was centrifuged at 12,000 g for 15 minutes at 4°C. The aqueous phase containing the RNA extracted from RNP was transferred to a new tube and mixed with 500 μl of 100% (v / v) ethanol. The sample was passed through an RNA Clean & Concentrator-5 spin column (Zymo Research). After washing the spin column with wash buffer, 13.6 units of DNAse (QIAGEN) were applied to the spin column and incubated at room temperature for 15 minutes to remove residual DNA. After washing the spin column three times with buffer, the bound RNA was eluted with 15 μl of nuclease-free water. Subsequently, 500 ng of purified RNA was used for RNA library preparation. RNA libraries were prepared using the SMARTer smRNA-Seq Kit for Illumina (TAKARA-BIO) according to the manufacturer's instructions. The resulting NGS adapter ligation cDNA was amplified for 8 cycles by PCR using full-length Illumina NGS indexing primers. Amplified DNA fragments of 200-500 bp were size-selected by 3% (w / v) agarose gel electrophoresis and purified by gel extraction using the FastGene Gel / PCR Extraction Kit (FastGene). The DNA fragments were pooled in equimolar ratios and the concentration was adjusted to 50 pM using the Qubit dsDNA HS Assay Kit and Qubit 3.0 Fluorometer (Thermo Fisher Scientific Inc.). The pooled libraries were mixed with 0.25 volume of 50 pM PhiX control v3 (Illumina Inc.) and then used for 2 × 150 bp paired-end sequencing on iSeq100 (Illumina Inc.). The reads were trimmed for the adapter and aligned to the template sequence using Bowtie2 software. The identified sequences of small RNAs bound to each TnpB-like RNA guide nuclease were defined as the sequences of the corresponding gRNAs (Figure 4, fused CDS and NGS reads). Information on the gRNAs is summarized in the table. [Table 9]

[0057] Example 5: Construction of a 7N-TAM library plasmid The TAM sequence of the TnpB-like RNA guide nuclease was determined using a plasmid library containing seven randomized nucleotides (7N). To construct the 7N-TAM plasmid library, ssDNA containing seven randomized nucleotides (5'-AGCTATGACCATGATTACGAATTCNNNNNNNCTGCAGGAGCAAAGACC-3' (SEQ ID NO: 87)) was converted to dsDNA using reverse transcriptase (ReverTra Ace, Takara Corporation) with a primer (5'-GTCTTTGCTCCTGCAG-3' (SEQ ID NO: 88)). This dsDNA fragment was assembled with pUC18 containing the NGS adapter sequence using the NEBuilder HiFi DNA Assembly Cloning Kit (NEB) to produce the 7N-TAM plasmid library (Table 10, plasmid sequences). E. coli DH5α cells were transformed with the reaction mixture and cultured at 37°C on LB plates supplemented with 100 g / ml ampicillin. Over 100,000 colonies were washed from the plate, and the plasmid library was extracted using the NucleoBond Xtra Midiprep kit (Takara Corporation). [Table 10]

[0058] Example 6: Screening of TAM sequences by in vitro transcription / translation In vitro transcription / translation (IVTT) reactions were performed using the PUREfrex2.0 reconstituted cell-free protein synthesis kit (GeneFrontier). Following the product instructions, DNA templates encoding a T7 promoter-driven TnpB-like RNA guide nuclease and a T7 promoter-driven gRNA scaffold with a spacer sequence (20nt:5'-GAATTCGTAATCATGGTCAT-3' (SEQ ID NO: 90)) targeting the TAM library were prepared by PCR using the TnpB-like RNA guide nuclease synthesis gene as the template. Illumina's Nextera HT v2 dual index primer was used as the PCR primer for DNA template preparation. The IVTT reaction was performed using 100 ng of the TnpB-like RNA guide nuclease template, 125 ng of the corresponding gRNA template with the spacer sequence, and 60 ng of the TAM library plasmid in 20-25 μl of reaction mixture. An IVTT reaction without the gRNA DNA template was used as a control. The reaction was incubated at 37°C for 4 hours, then quenched with 5 μl of 100 mg / ml RNase A (Nippon Gene) and incubated at 37°C for 30 minutes. After adding 5 μl of 1% SDS, the reaction was incubated at 50°C for 10 minutes to denature the protein. Four units of protease K (FUJIFILM) were added, and the protein was digested by incubation at 37°C for 90 minutes. The undigested TAM library plasmid was recovered from the reaction product using the FastGene Gel / PCR Extraction Kit and used as a template for subsequent PCR. A DNA fragment of approximately 250 bp containing the TAM motif and target sequence was amplified by 30 cycles of PCR using KOD One PCR Master Mix (Toyobo) and Illumina full-length Nextera NGS indexing primers. The amplified DNA fragments were separated by 2.5% (w / v) agarose gel electrophoresis and then purified by gel extraction. The preferentially depleted TAM motifs were identified by NGS amplicon sequencing of TAM library plasmids, as described above. TAM priority was characterized from FASTQ files using a custom Python script.In short, 7-nucleotide TAMs were extracted and counted. TAM frequencies were normalized for the sequencing depth of each sample. Sequences were generated using WebLogo version 2.8.2 (http: / / weblogo.berkeley.edu / ) with TAMs showing at least a 10-fold drop compared to the control (Figure 5).

[0059] Example 7: In vitro analysis of dsDNA cleavage The 7N sequence was converted to a TAM sequence by site-directed mutagenesis, and a substrate plasmid containing the TAM sequence was synthesized. Substrate plasmids containing different target sequences (5'-GTTCTCCAGGCTGCTATCCTTAGCA-3' (SEQ ID NO: 91), 5'-GAATTCGTAATCATGGTCATAGCTG-3' (SEQ ID NO: 92)) were synthesized by site-directed mutagenesis. A 716 bp substrate dsDNA was amplified by PCR using the TAM sequence plasmid and primer 1 (5'-AAAGGGGATGTGCTGCAAGG-3' (SEQ ID NO: 93)) and primer 2 (5'-TATCTTTATAGTCCTGTCGG-3' (SEQ ID NO: 94)). The amplified dsDNA fragment (20 nM) was mixed with purified TnpB-like RNA guide nuclease RNP (1, 2, 4 μM) corresponding to the target sequence in 5 μl of buffer D and incubated for 2 hours. The reaction mixture was quenched with 2 μl of Proteinase K (Fujifilm Wako Pure Chemical Industries) and incubated at 60°C for 5 minutes. The reaction product was then subjected to MultiNA microchip electrophoresis using a system (Shimadzu Inc.) (Figure 6).

[0060] Example 8: In vitro plasmid cleavage analysis The 7N sequence was converted to a TAM sequence by site-directed mutagenesis, and substrate plasmids containing the TAM sequence were synthesized. Specifically, for T5, substrate plasmids were constructed containing either 5'-TTAT-3' or 5'-AACAT-3', and for T8-2, either 5'-TTCAT-3' or 5'-AACAT-3', attached to the 5' side of the target sequence 5'-GAATTCGTAATCATGGTCATAGCTG-3' (SEQ ID NO: 92). 50 ng of the substrate plasmid was mixed with 5 μl of Buffer D and purified TnpB-like RNA guide nuclease RNP (4 μM) corresponding to the target sequence, and incubated at 37°C for 2 hours. The reaction mixture was quenched with 2 μl of Proteinase K (Fujifilm Wako Pure Chemical Industries) and incubated at 60°C for 5 minutes. The reaction product was analyzed on a 1.0% (w / v) agarose gel (Figure 7). When the TAM sequence corresponds to the correct one, a change in mobility due to cleavage was observed.

[0061] Example 9: In vitro plasmid cleavage analysis (confirmation of temperature dependence) The 50 ng substrate plasmid used in Example 8 was mixed with purified TnpB-like RNA guide nuclease RNP (4 μM) corresponding to the target sequence in 5 μl of buffer D. The reaction temperature was varied in 5 °C increments from 20 °C to 50 °C, and the mixture was incubated for 2 hours under each condition. The reaction mixture was quenched with 2 μl of Proteinase K (Fujifilm Wako Pure Chemical Industries) and incubated at 60 °C for 5 minutes. The reaction product was analyzed on a 1.0% (w / v) agarose gel (Figure 8). This result indicates that the TnpB RNA guide nuclease used in this invention has a lower activity temperature (around 37 °C) compared to typical TnpB proteins and TnpB-derived genes, which have maximum activity around 40-50 °C.

[0062] Example 10: Construction of T8 protein and gRNA expression vector Plasmid vectors for expressing T8-1 protein or T8-2 protein and gRNA in human cultured cells were constructed as follows. The plasmid vector for expressing the T protein is a pUC-based transient expression plasmid vector (SEQ ID NO: 95) that contains a CAG promoter, has a multicloning site (MCS) at the C-terminus of the T protein to which SV40NLS and NuleoplasminNLS are conjugated, and has a bGH polyA signal. The human codon-optimized T8-1 CDS sequence of the T protein (SEQ ID NO: 96, synthesized by IDT) or the human codon-optimized T8-2 sequence (SEQ ID NO: 97, synthesized by IDT) were subcloned into the aforementioned MCS, respectively. The plasmid vector for expressing gRNA is a pUC-based transient expression plasmid vector (SEQ ID NO: 98) that contains a human U6 promoter, a non-coding RNA cloning site, and a hammerhead ribozyme sequence at its 3' end. First, the gT8-1 gRNA scaffold sequence (SEQ ID NO: 99, synthesized by IDT) or the T8-2 gRNA scaffold sequence (SEQ ID NO: 100, synthesized by IDT) was introduced into the aforementioned gRNA expression vector. Furthermore, six different sequences targeting the hAAVS1 gene (SEQ ID NOs: 101-106) were subcloned.

[0063] Example 11: Evaluation of genome editing efficiency using HEK293FT Each solution containing 100 ng of the two types of T8 protein expression vectors (expressing T8-1 or T8-2) constructed in Example 10 and 50 ng of gRNA expression vectors (one of each of six gRNA expression vectors targeting the hAAVS1 gene) was mixed with Lipofectamine 3000 reagent, and 96-well cultured HEK293FT cells (~10 4 The product was added to cells (per well), and transfection was performed using all 12 combinations. For comparison, a control was prepared using a GFP protein expression plasmid (pEGFP-N1, Clonetech). After culturing for 3 days, the DNA of HEK293FT cells was extracted with alkaline buffer (0.1N NaOH), and 100 pg / μL of crudely purified DNA was used as a template. Using KOD One PCR master Mix and the first round PCR primer set (sequences 110 and 111 if gRNA1, gRNA2, or gRNA3 was used, and sequence 112 and 113 if gRNA4, gRNA5, or gRNA6 was used), a partial sequence of the hAAVS1 gene containing the genome editing target sequence was subjected to PCR under the following conditions to amplify the target sequence. Similarly, using a 20-fold dilution of the first-round PCR solution as a template, a PCR reaction was performed under the following conditions using the second-round PCR primer set and second-round PCR primers containing the Combinatorial Dual Index (CDI) sequence (5′-AATGATACGGCGACCACCGAGATCTACACNNNNNNNNTCGTCGGCAGCGTC-3′ (SEQ ID NO: 107) and 5′-CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTCTCGTGGGCTCGG-3′ (SEQ ID NO: 108)). PCR amplification products with adapter sequences and index sequences at both ends for sequencing were obtained using an Illumina sequencing instrument. PCR conditions: 98°C / 2 minutes, 98°C / 10 seconds, 55°C / 5 seconds, 68°C / 3 seconds, repeated 35 times. (Applicable to both Round 1 and Round 2) The obtained PCR products were subjected to agarose gel electrophoresis, and the target nucleotide sequence portion was excised and purified. The purified DNA was concentrated using the Qubit dsDNA HS Assay Kit and Qubit 3.0 Fluorometer (Thermo Fisher Scientific, Waltham, MA, USA) and adjusted to 50 pM. The pooled library was mixed with 0.25 volume of 50 pM PhiX control v3 (Illumina) and then used for 2 × 150-bp paired-end sequencing on iSeq100 (Illumina, San Diego, CA, USA). The obtained Fastq files were analyzed using CRISPResso2 (Clement et al., 2019 Nature Biotechnology). When the T8 protein exhibits a cleavage pattern similar to TnpB, it is considered to be an insertion or deletion within 10 nucleotides downstream of the target sequence, or genome editing induced by the T8 protein; therefore, the quantitative window width for mutations was set to 10 bp increments from the expected cleavage site (Figures 9 and 10). For "genome editing efficiency," we used the percentage of all sequenced reads that had insertions or deletions within the quantitative window width. Figure 9 shows the results of evaluating the genome editing efficiency of each gRNA1-gRNA6 under conditions where they were co-expressed simultaneously with T8-1 or T8-2, respectively, using the procedure described above. The vertical axis represents genome editing efficiency (the percentage of reads with insertion or deletion mutations in the target region out of all reads), and the horizontal axis represents the conditions of the transformed vector combination. While the genome editing efficiency was almost 0 in the control, in both the T8-1 and T8-2 cases, the highest genome editing efficiency (approximately 0.006) was observed when co-expressed with gRNA3. Next, gRNA5 was efficient (approximately 0.0015). Although T8-1 and T8-2 showed similar genome editing efficiencies, gRNA1, gRNA2, and gRNA4 showed higher genome editing efficiency when co-expressed with T8-2 compared to when co-expressed with T8-1. Figure 10 shows a specific example of genome editing when a T8-2 protein expression vector and a gRNA3 expression vector were simultaneously introduced into HEK293FT cells and the genome sequence was analyzed. The target genome sequence is shown as the Reference, the expected gRNA binding site is indicated by the sgRNA (rectangle), and the expected cleavage site is indicated by a dashed line. Different sequences output by the sequencer were aligned, and the number of reads and the percentage of total reads are shown. Base substitutions are shown in bold, and deletions are shown with "-". When genome editing occurs, insertions or deletions are observed near the expected cleavage site (the dashed line), and several types of such reads (especially deletions of 3-8 bases) were detected at a rate of about 0.1%. From the above, it was found that in HEK293FT cells co-introduced with a T8-2 expression vector and a gRNA3 expression vector, genome editing mainly consisting of deletions occurs in the target genome region.

[0064] Example 12: Construction of a modified T8-2 guide RNA expression vector Plasmid vectors for expressing T8-2 guide RNA in human cultured cells were constructed as follows. The plasmid vector for expressing the gRNA is a pUC-based transient expression plasmid vector (SEQ ID NO: 98) that contains a human U6 promoter, a non-coding RNA cloning site, and a hammerhead ribozyme sequence at its 3' end. First, a 152nt T8-2 guide RNA scaffold sequence (SEQ ID NO: 100, synthesized by IDT) was introduced into the gRNA expression vector. Furthermore, a sequence targeting the hAAVS1 gene (SEQ ID NO: 103) was subcloned. The guide RNA, composed of this T8-2 guide RNA scaffold sequence, target sequence, and hammerhead ribozyme sequence, was used as the gRNA original (SEQ ID NO: 114).

[0065] The modified T8-2 guide RNA expression vector was prepared as follows: A guide RNA expression vector was created by removing only the hammerhead ribozyme sequence from the transient expression plasmid vector (SEQ ID NO: 98) (SEQ ID NO: 115). A guide RNA sequence consisting only of the T8-2 guide RNA scaffold sequence (SEQ ID NO: 98) and the hAAVS1 gene target sequence (SEQ ID NO: 100) was cloned into the cloning site of this vector. This guide RNA sequence was designated as ddHDV (SEQ ID NO: 116). Next, four shortened guide RNAs (synthesized by IDT) were created by stepwise deleting nucleotide sequences from the 5' end of the guide RNA scaffold sequence within the ddHDV sequence: gRNA1 (guide RNA scaffold sequence length 131 nt, SEQ ID NO: 117), gRNA2 (guide RNA scaffold sequence length 111 nt, SEQ ID NO: 118), gRNA3 (RNA scaffold sequence length 100 nt, SEQ ID NO: 119), and gRNA4 (guide RNA scaffold sequence length 77 nt, SEQ ID NO: 120). These were then cloned to create a total of five modified T8-2 guide RNA expression vectors.

[0066] Example 13: Evaluation of genome editing efficiency using modified T8-2 guide RNA with HEK293FT Using the T8-2 protein expression vector constructed in Example 10 and the modified T8-2 guide RNA expression vector constructed in Example 12, the effect of modifying the guide RNA scaffold sequence on genome editing efficiency was evaluated using the method described in Example 11 (Figure 11). For comparison, a condition with a GFP protein expression plasmid (pEGFP-N1, Clonetech) was prepared (referred to as Control). In experiments where a T8-2 protein expression vector and gRNA expression vectors with the same target sequence but different gRNA coding regions were co-introduced, the genome editing efficiency of the hAAVS1 gene was compared. No base insertions or deletions were observed in the hAAVS1 gene in the control group. When gRNA original, which has a full-length 152nt T8-2 guide RNA scaffold sequence, or ddHDV, which has only the HDV sequence removed, were co-expressed with T8-2, the genome editing efficiency was approximately 0.7%. This result indicates that the removal of the hammerhead ribozyme sequence in the guide RNA does not affect genome editing efficiency. While the genome editing efficiency was also approximately 0.7% when using gRNA1, which had the T8-2 guide RNA scaffold sequence shortened to 131nt, it improved to approximately 1.1% when using gRNA2, which had the guide RNA scaffold sequence shortened to 111nt, and to approximately 0.86% when using gRNA3, which had it shortened to 100nt. However, no genome editing activity was observed with gRNA4 shortened to 77nt. These results indicate that deleting 21nt from the 5' end of the original T8-2 guide RNA scaffold sequence improves genome editing activity, and further, that the RNA sequences contained in the gRNA scaffolds of gRNA3 and gRNA4 play an important role in maintaining the DNA cleavage activity of T8-2.

[0067] Example 14: Construction of a T8-2 mutant protein expression vector Similar to Example 11, plasmid vectors expressing 24 different human codon-optimized T8-2 mutant proteins (SEQ ID NOs. 121–168) were constructed instead of the T8-2 protein.

[0068] Example 15: Evaluation of genome editing efficiency using T8-2 mutant protein with HEK293FT Similar to Example 12, the genome editing efficiency of the T8-2 mutant protein was evaluated. The plasmid vector used to express the gRNA was the one containing Sequence ID No. 118, which showed the highest editing efficiency among the six gRNA vectors used in Example 13 (Figure 12). Figure 12 shows the results of evaluating genome editing efficiency using the procedure described above, under conditions where either the T8-2 protein without mutations or each of the 24 T8-2 mutant proteins were simultaneously co-expressed with a gRNA3 expression vector. The vertical axis represents genome editing efficiency (percentage of all reads that had insertion or deletion mutations in the target region), and the horizontal axis represents the conditions of the transformed vector. While the genome editing efficiency was almost 0 in the control and about 0.3% when the non-mutated T8-2 protein was expressed, it showed a genome editing efficiency of about 1% when the T8-2 mutant protein had the K89R mutation, the Q284K mutation, or the R291K mutation, respectively. Furthermore, in all cases, the T8-2 mutant protein with both the K89R and Q284K mutations, or the K89R and R291K mutations, or the R291K and Q284K mutations, showed a genome editing efficiency of about 2%. Furthermore, T8-2 mutant proteins possessing the K89R, R291K, and Q284K mutations simultaneously showed a genome editing efficiency of approximately 2.5%. Therefore, it was concluded that the K89R, R291K, and Q284K mutations each increase the genome editing efficiency of the T8-2 protein, and that combinations of these mutations can yield T8-2 mutant proteins with enhanced genome editing capabilities.

[0069] Example 16: Construction of a fusion protein expression plasmid and gRNA expression vector composed of inactive T8-2 and the catalytic domain of an epigenetic factor histone acetyltransferase. In known TnpB proteins, a typical DED active site sequence is known to be contained within the endonuclease domain RuvC. It is known that substituting the N-terminal aspartic acid with alanine causes the DNA cleavage activity to be lost. Furthermore, it is known that expressing a fusion protein composed of a known inactive Cas9 and the catalytic domain (P300 core) of a histone acetyltransferase, along with a guide RNA targeting the promoter sequence, can promote acetylation of histone proteins near the target sequence, thereby inducing transcriptional activation of nearby genes through changes in the epigenetic state of the promoter region. With the aim of utilizing T8-2 in gene transcription activation technology, we aimed to create an artificial transcription activator composed of an inactive T8-2 and a P300 core, which are created by amino acid substitution of the DED active site sequence. A plasmid vector for expressing a fusion protein and gRNA composed of inactive T8-2 and the histone acetyltransferase catalytic domain P300core in human cultured cells was constructed as follows. A human codon-optimized CDS sequence (SEQ ID NO: 169) encoding a fusion protein (dead T8-2-P300 core) was created (SEQ ID NO: 170, synthesized by IDT). This sequence consists of an inactive T8-2 sequence in which the 308th aspartic acid residue in the DED active site sequence of T8-2 is replaced with alanine, a hemagglutinin tag at the C-terminus of the inactive T8-2, two SV40 NLSs, and a P300 core sequence of human-derived histone acetyltransferase. The fusion protein expression vector was created by cloning the CDS sequence between the CBh promoter and the bGHpolyA signaling pathway of a transient expression plasmid vector (SEQ ID NO: 95). Furthermore, vectors expressing T8-2 guide RNA were created by subcloning either a DNA sequence (SEQ ID NO: 171) targeting the Tet operator sequence (TetO) on the 3' side of the guide RNA scaffold sequence of the T8-2 guide RNA expression vector created in Example 10, or a DNA sequence (SEQ ID NO: 172) targeting a sequence not present in the reporter vector or the human genome.

[0070] Example 18: Construction of an eGFP expression reporter vector to evaluate transcriptional activation by a dead T8-2-P300 core. The following eGFP expression reporter vector was constructed to evaluate transcriptional activation by a dead T8-2-P300 core in human cultured cells. The reporter vector is a transient expression plasmid vector based on pUC, which has a TRE3G promoter with a repeat sequence consisting of seven Tet operator sequences (TetO), and a Kozak sequence, an eGFP CDS sequence, and an SV40 polyA signal at its 3' end (SEQ ID NO: 173, synthesized by VectorBuilder).

[0071] Example 19: Evaluation of transcriptional activation ability of a dead T8-2-P300 core using HEK293FT 300 ng of the dead T8-2-P300 core expression vector constructed in Example 17, a T8-2 guide RNA expression vector (sgRNA (TetO target)) having the TetO sequence of the eGFP reporter vector as its target sequence, or a guide RNA expression vector (sgRNA (No target)) having a reporter vector and a sequence not present in the human genome as its target sequence, were mixed with 300 ng of the eGFP expression reporter vector prepared in Example 18, along with Lipofectamine 3000 reagent. These mixtures were then used to culture HEK293FT cells in 96 wells (~10 4 The solution was added to the cells (per well) and transfection was performed. Transfected HEK293FT cells were cultured in Dulbecco's modified Eagle medium containing 10% fetal bovine serum at 37°C under 5% CO2 conditions. After 24 hours of culture, the expression level of eGFP was measured by fluorescence microscopy, and the transcriptional activation ability was evaluated (Figure 13). The cumulative eGFP fluorescence intensity of cultured cells introduced with the dead T8-2-P300 core expression vector and gRNA (TetO) was approximately 7.1 times higher than that of cells introduced with gRNA (No target). This result indicates that the dead T8-2-P300 core fusion protein induces transcriptional activation of genes located near the target sequence by altering the epigenetic state.

Claims

1. (i) A TnpB-like RNA guide nuclease comprising the amino acid sequence of SEQ ID NO: 20 or 22, or a mutant protein thereof, wherein the mutant protein has at least 90% sequence identity with the amino acid sequence and has TnpB-like RNA guide nuclease activity, (ii) A guide RNA comprising a guide sequence and a sequence containing a gRNA scaffold sequence, Here, when using a TnpB-like RNA guide nuclease consisting of the amino acid sequence of SEQ ID NO: 20, or a mutant protein thereof, the gRNA scaffold sequence includes the nucleic acid sequence of SEQ ID NO: 99, or a sequence having at least 90% identity with said sequence. A TnpB-like RNA guide nuclease complex in which, when using a TnpB-like RNA guide nuclease consisting of the amino acid sequence of SEQ ID NO: 22, or a mutant protein thereof, the gRNA scaffold sequence includes the nucleic acid sequence of SEQ ID NO: 100, or a sequence having at least 90% identity with said sequence.

2. The TnpB-like RNA guide nuclease complex according to claim 1, wherein the TAM sequence of the TnpB-like RNA guide nuclease is a sequence containing TCAT.

3. The TnpB-like RNA guide nuclease complex according to claim 2, wherein the TAM sequence comprises TTCAT.

4. The TnpB-like RNA guide nuclease complex according to claim 1, wherein the mutant protein has an amino acid substitution at at least one position selected from the group consisting of Q at position 284, K at position 89, and R at position 291.

5. In the aforementioned mutant protein, If there is an amino acid substitution at position 284, it is either Q284R or Q284K. If there is an amino acid substitution at position 89, it is K89R. If there is an amino acid substitution at position 291, it is R291K. The TnpB-like RNA guide nuclease complex according to claim 4.

6. The TnpB-like RNA guide nuclease complex according to claim 1, wherein the TnpB-like RNA guide nuclease is a TnpB-like RNA guide nuclease consisting of the amino acid sequence of SEQ ID NO: 22, or a mutant protein thereof, and the gRNA scaffold sequence comprises a nucleic acid consisting of SEQ ID NO: 117 (gRNA1), SEQ ID NO: 118 (gRNA2), SEQ ID NO: 119 (gRNA3), or a nucleotide sequence having 90% sequence identity to said sequence.

7. A method for cleaving target DNA using a TnpB-like RNA guide nuclease complex according to any one of claims 1 to 6 (excluding methods for cleaving in living human cells).

8. The aforementioned cutting method, A nucleic acid comprising an expression cassette encoding the aforementioned TnpB-like RNA guide nuclease, The cleavage method according to claim 7, comprising introducing a nucleic acid containing an expression cassette encoding the guide RNA into a cell.

9. The cutting method according to claim 7, wherein the cutting method is used as a genome editing method including sequence-specific knockout and sequence-specific knock-in.

10. A genome editing kit comprising the TnpB-like RNA guide nuclease complex described in any one of claims 1 to 6.

11. The genome editing kit according to claim 10, wherein the guide RNA containing the gRNA scaffold sequence contains a guide sequence designed from the cleavage target sequence.

12. The genome editing kit according to claim 10, wherein the TAM sequence of the TnpB-like RNA guide nuclease is a sequence containing TCAT.

13. The genome editing kit according to claim 10, wherein the genome editing kit is used for sequence-specific knockout and sequence-specific knock-in.