TnpB-like RNA-guided nuclease complex

A novel TnpB-like RNA-guided nuclease complex addresses the limitations of existing genome editing technologies by recognizing a TAM sequence, enabling precise and efficient DNA cleavage and transcription regulation across diverse organisms.

JP7821459B2Active Publication Date: 2026-02-27NATIONAL INSTITUTE OF ADVANCED INDUSTRIAL SCIENCE & TECHNOLOGY +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025555860
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2024-12-19
Publication Date
2026-02-27
Estimated Expiration
2044-12-19

AI Technical Summary

Technical Problem

Existing genome editing technologies, such as CRISPR/Cas9 and CRISPR/Cas12a, are limited by the requirement for specific PAM sequences, which can vary depending on the type of Cas protein, making it challenging to target desired DNA sequences efficiently.

Method used

Development of a novel TnpB-like RNA-guided nuclease complex comprising a TnpB-like RNA-guided nuclease with a specific amino acid sequence and a guide RNA, capable of recognizing a 3-7 base TAM sequence and cleaving target DNA, allowing for more flexible and efficient genome editing.

Benefits of technology

The TnpB-like RNA-guided nuclease complex enables precise and efficient cleavage of target nucleic acids, facilitating genome editing and transcription regulation, with applications in various organisms including plants, algae, and microorganisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007821459000011
    Figure 0007821459000011
  • Figure 0007821459000012
    Figure 0007821459000012
  • Figure 0007821459000013
    Figure 0007821459000013
Patent Text Reader

Abstract

The purpose of the present invention is to provide: a new RNA guide nuclease that can be applied to genome editing technique; and use of the same. Discovered is a TnpB-like RNA guide nuclease that functions in an RNA guide endonuclease complex that can be applied to genome editing technique. As a result, provided is a TnpB-like RNA guide nuclease complex comprising a TnpB-like RNA guide nuclease and a gRNA having a guide sequence and a gRNA scaffold sequence.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a novel TnpB-like RNA-guided nuclease complex, a technique for cleaving target DNA using the complex, a genome editing technique involving cleavage, and a technique for regulating transcription in target DNA using the TnpB-like RNA-guided nuclease complex. [Background technology]

[0002] Genome editing is a technique for introducing (editing) mutations at targeted locations in the genome of a target organism. Genome editing using CRISPR-Cas is based on the discovery of an adaptive immune system in bacteria and archaea against foreign viruses and plasmids. It utilizes Cas, an RNA-guided endonuclease involved in this adaptive immune system. DNA from foreign viruses or plasmids is incorporated into the CRISPR locus by the action of Cas proteins. Once incorporated into the CRISPR locus, a portion of the foreign DNA functions as a template for producing crRNA. The produced crRNA forms a complex with tracrRNA, functioning as a guide strand to recruit Cas proteins to foreign DNA complementary to the crRNA, which then cleaves the foreign DNA near the PAM sequence, eliminating the foreign DNA and exerting its immune function. Focusing on the ability to cleave a specific sequence, guide RNAs that function as crRNA and tracrRNA can be designed to complement the target DNA sequence, allowing DNA cleavage at the desired location. Therefore, compared to existing mutagenesis methods such as radiation, the advantage of this method is its efficiency in obtaining the desired strain from an overwhelmingly smaller number of mutants, and it is used for a variety of purposes.

[0003] In typical genome editing technologies such as CRISPR / Cas9 and CRISPR / Cas12a, the target sequence is primarily determined by a region called the spacer sequence of the gRNA or crRNA that makes up the gene. The gRNA or crRNA spacer sequence typically serves as a target sequence of 20–24 bases, which can be freely modified. In addition to this target sequence, the nucleases Cas9 and Cas12a require a PAM sequence, a 2–4 ​​base sequence, in the target nucleic acid sequence for their function. Therefore, genome editing is achieved by specifically recognizing a total of approximately 22–28 bases, consisting of the PAM sequence and the target sequence. The target DNA sequence must contain a PAM sequence, but PAM sequences vary depending on the type of Cas protein family. Genome editing of the desired sequence is possible by selecting a Cas protein that matches the target DNA sequence.

[0004] It is speculated that the Cas protein, an RNA-guided endonuclease, originates from the IscB and TnpB proteins contained in transposons (Non-Patent Document 1: J. Bacteriol. 198, 797-807 (2015)). The IscB and TnpB proteins, nucleases contained in transposons, have been reported to function as RNA-guided endonucleases (Non-Patent Document 2: Science 374, 57-65 (2021), Non-Patent Document 3: Nature 599, 692-696 (2021)). It has also been reported that a complex containing a protein classified as TnpB, which contains a Ruv-C nuclease domain, and an ωRNA molecule can be used for genome editing (Patent Document 1: International Publication No. 2022 / 159892). Active research is being conducted into functional analysis of TnpB and the identification of TnpB orthologs and their use as new genome editing tools (Non-Patent Document 4: The CRISPR Journal. Jun 2023 pp. 232-242; Non-Patent Document 5: Nat Biotechnol., 2023); Non-Patent Document 6: Nature 620, 660-668, 2023; Non-Patent Document 7: Science Advances Sep 2023 vol. 9 Issue 39; Non-Patent Document 8: Nucleic Acids Research, gkad1053, 2023; Non-Patent Document 9: Proc Natl Acad Sci USA. Nov 2023). 28;120(48):e2308224120.). [Prior art documents] [Patent documents]

[0005] [Patent Document 1] International Publication No. 2022 / 159892 [Non-patent literature]

[0006] [Non-Patent Document 1] J. Bacteriol. (2015) 198, 797-807 [Non-patent document 2] Science (2021)374, 57-65 [Non-patent document 3] Nature (2021) 599, 692-696 [Non-patent document 4] The CRISPR Journal.(2023) vol. 6, No. 3, p.232-242 [Non-Patent Document 5] Nat Biotechnol., 2023, June 29, p.1-13 [Non-patent document 6] Nature, 2023,620,660-668 [Non-Patent Document 7] Science Advances Sep 2023 vol. 9 Issue 39, eadlk0171 [Non-patent document 8] Nucleic Acids Research, gkad1053,2023 [Non-Patent Document 9] Proc Natl Acad Sci USA. 2023 Nov 28;120(48):e2308224120. [Non-Patent Document 10] Int J Mol Sci., 2020 Apr 25;21(9):3038. doi: 10.3390 / ijms21093038. Summary of the Invention [Problem to be solved by the invention]

[0007] The aim is to identify and utilize novel RNA-guided nucleases that can be applied to genome editing technology. [Means for solving the problem]

[0008] The present inventors conducted extensive research with the aim of identifying new proteins not annotated with TnpB that form RNA-guided endonuclease (hereinafter also referred to as RGN) complexes applicable to genome editing technology. As a result, they discovered a TnpB-like RNA-guided nuclease that is not classified as TnpB, leading to the present invention. The present invention therefore relates to: [1] (i) the following: YKGRTFNKMINNGSKGQYNxR-SxNxLKWRG (SEQ ID NO: 109) wherein x represents any amino acid. a TnpB-like RNA-guided nuclease having a length of 300 to 700 amino acids, (ii) A TnpB-like RNA-guided nuclease complex comprising a guide sequence and a guide RNA comprising a sequence including a gRNA scaffold sequence. [2] The TnpB-like RNA-guided nuclease complex according to Item 1, wherein the TnpB-like RNA-guided nuclease is a protein consisting of an amino acid sequence selected from the group consisting of SEQ ID NOs: 16, 20, 22, 24, 26, 40, 50, 52, 54, and 56. [3] The TnpB-like RNA-guided nuclease complex according to Item 1, wherein the TnpB-like RNA-guided nuclease is a mutant protein having a mutation involving an amino acid substitution with respect to the protein according to Item 2, and has RNA-guided nuclease activity. [4] The TnpB-like RNA-guided nuclease complex according to Item 3, wherein the mutant protein has an amino acid sequence that has at least 75% sequence identity to the amino acid sequence of the TnpB-like RNA-guided nuclease complex according to Item 2. [5] The TnpB-like RNA-guided nuclease complex according to Item 3, wherein the mutant protein has structural homology to the original protein in the RuvC domain, and the structural homology is such that, when the original protein and the mutant protein are structurally aligned using PyMOL software, the C root mean square deviation (RMSD) between the backbone of the original protein and the backbone of the mutant protein is 1.5 or less and the TM value is 0.9 or more. [6] The TnpB-like RNA-guided nuclease complex according to any one of items 1 to 5, wherein the TAM sequence of the TnpB-like RNA-guided nuclease is a sequence containing TCAT. [7] The TnpB-like RNA-guided nuclease complex according to Item 6, wherein the TAM sequence comprises TTCAT. [8] The TnpB-like RNA-guided nuclease complex according to any one of Items 1 to 7, wherein the TnpB-like RNA-guided nuclease is a TnpB-like RNA-guided nuclease consisting of the amino acid sequence of SEQ ID NO: 22 or a mutant protein thereof, and the gRNA scaffold sequence comprises the sequence of SEQ ID NO: 58 or a sequence having at least 70% identity to said sequence. [9] The TnpB-like RNA-guided nuclease complex according to Item 8, wherein the mutant protein has an amino acid substitution at at least one position selected from the group consisting of Q at position 284, R at position 458, K at position 89, S at position 72 and / or S at position 75, R at position 291, S at position 201, L at position 278, and S at position 315.

[10] The mutant protein, If there is an amino acid substitution at position 284, it is Q284R or Q284K; If there is an amino acid substitution at position 89, it is K89R; when there is an amino acid substitution at position 72 or 75, it is S72H, S72N, S72Q, S72K, S72R, S75H, S75N, S75Q, S75K, or S75R; If there is an amino acid substitution at position 291, it is R291E or R291K; If there is an amino acid substitution at position 201, it is Q201H or Q201S; If there is an amino acid substitution at position 278, it is L278K or L278R; If there is an amino acid substitution at position 315, it is S315A or S135V. 10. The TnpB-like RNA-guided nuclease complex according to item 9.

[11] The TnpB-like RNA-guided nuclease complex according to any one of Items 1 to 10, wherein the TnpB-like RNA-guided nuclease is a TnpB-like RNA-guided nuclease consisting of the amino acid sequence of SEQ ID NO: 22 or a mutant protein thereof, and the gRNA scaffold sequence comprises a nucleic acid consisting of SEQ ID NO: 117 (gRNA1), SEQ ID NO: 118 (gRNA2), SEQ ID NO: 119 (gRNA3), or a nucleotide sequence having 90% sequence identity to said sequences.

[12] A method for cleaving a target DNA using the TnpB-like RNA-guided nuclease complex described in any one of items 1 to 11.

[13] The cutting method comprises: a nucleic acid comprising an expression cassette encoding the TnpB-like RNA-guided nuclease; Item 13. The cleavage method according to Item 12, comprising introducing into a cell a nucleic acid comprising an expression cassette encoding the guide RNA.

[14] The cleavage method described in Item 12, wherein the cleavage method is used as a genome editing method including sequence-specific knockout and sequence-specific knock-in.

[15] (i) a TnpB-like RNA-guided nuclease of the TnpB-like RNA-guided nuclease complex according to any one of items 1 to 11, which has an inactivating mutation and is fused with a transcriptional regulatory domain; (ii) A TnpB-like RNA-guided nuclease complex containing an inactivating mutation, the TnpB-like RNA-guided nuclease complex comprising a guide sequence and a guide RNA comprising a sequence including a gRNA scaffold sequence.

[16] A method for regulating transcription of a target DNA using the TnpB-like RNA-guided nuclease complex containing an inactivating mutation according to Item 15.

[17] The method for regulating transcription comprises: an expression cassette encoding the TnpB-like RNA-guided nuclease containing the inactivating mutation; Item 16. The method for regulating transcription of a target DNA according to Item 15, comprising introducing into a cell an expression cassette encoding the guide RNA.

[18] A genome editing kit comprising the TnpB-like RNA-guided nuclease complex according to any one of items 1 to 11.

[19] The genome editing kit according to Item 18, wherein the guide RNA containing the gRNA scaffold sequence contains a guide sequence designed from the cleavage target sequence.

[20] The genome editing kit according to Item 18, wherein the TAM sequence of the TnpB-like RNA-guided nuclease is a sequence containing TCAT.

[21] The genome editing kit according to any one of Aspects 18 to 20, wherein the genome editing kit is used for sequence-specific knockout and sequence-specific knockin. [Effects of the Invention]

[0009] According to the present invention, a novel RNA-guided endonuclease enables cleavage of target nucleic acids, and can be applied to genome editing technology. [Brief explanation of the drawings]

[0010] [Figure 1] Figure 1 shows the phylogenetic relationships of the five T protein groups (clades 1 to 5). The phylogenetic tree was constructed by the neighbor-joining method using the protein sequences listed in Table 1. [Figure 2] Figure 2 shows the 3'-terminal sequences of the genes encoding TnpB-like RNA-guided nucleases belonging to each clade. We focused on the 3'-terminal boundary of the region of high homology between each clade, and determined this boundary as the TEM sequence. [Figure 3] Figure 3 shows the results of SDS-PAGE and Coomassie blue staining of the TnpB-like RNA-guided nuclease solution obtained by expressing and column-purifying the recombinant His-MBP-TEV fusion nuclease in E. coli (A: T2, B: T3, C: T4, D: T5, E: T6, F: T7, G: T8-1, H: T8-2). [Figure 4] Figure 4 shows the results of next-generation sequencing analysis of the RNA complexed with TnpB-like RNA-guided nuclease. The RNA complexed with TnpB-like RNA-guided nuclease was shown to partially overlap with the open reading frame of the TnpB-like RNA-guided nuclease gene. [Figure 5] Figure 5 shows the results of analyzing the cleaved plasmids after applying a complex of TnpB-like RNA-guided nuclease and guide RNA to a plasmid library containing random TAMs and target sequences. Plasmids containing TCAT (T3), TTAT (T5 and T7), and TTCAT (T6, T8-1, and T8-2) were cleaved from the random sequences, resulting in the TAM sequences. [Figure 6] Figure 6 shows the results of in vitro cleavage of double-stranded DNA containing different target sequences amplified from purified plasmids containing TAM sequences, in which a complex of a guide RNA corresponding to the target sequence and a TnpB-like RNA-guided nuclease was reacted with the DNA. [Figure 7] Figure 7 shows the results of in vitro cleavage assays of purified plasmids containing different TAM sequences and different target sequences, in which a complex of guide RNA corresponding to the target sequence and TnpB-like RNA-guided nucleases (T5 and T8-2) was applied. The types of RNPs and TAM sequences used are shown at the top. No RNP is the negative control. [Figure 8] Figure 8 shows the results of in vitro cleavage activity of purified plasmids containing TAM sequences and target sequences for T5 and T8, treated with a complex of a guide RNA corresponding to the target sequence and a TnpB-like RNA-guided nuclease (T5 or T8-2) at various temperatures. [Figure 9]Figure 9 is a graph showing the genome editing efficiency when HEK293FT was transfected with an expression vector encoding the T8-1 protein or T8-2 and an expression vector encoding a guide RNA. The vertical axis shows the absolute value of genome editing efficiency, and the horizontal axis shows the conditions for transformation with T8-1 gRNA1-5 and T8-2 gRNA1-5, respectively. Control shows the genome editing efficiency when no T8 expression or guide DNA expression plasmid was included. [Figure 10] Figure 10 shows the genomic nucleic acid sequence (5' left, 3' right) near the cleavage site when HEK293FT is transfected with an expression vector encoding the T8-1 protein and an expression vector encoding a guide RNA. The sgRNA represents the target sequence and recognizes the complementary strand of the indicated DNA. The TAM sequence is located 3' to the target sequence and is not included in this figure. The dashed line indicates the predicted location of double-stranded cleavage. [Figure 11] Figure 11 shows the editing efficiency when modified T8-2 is introduced into the following guide RNAs: gRNA original, which has a T8-2 guide RNA scaffold sequence of 152 nt in total; ddHDV, which has only the HDV sequence removed; gRNA1, which has the T8-2 guide RNA scaffold sequence shortened to 131 nt; gRNA2, which has the guide RNA scaffold sequence shortened to 111 nt; gRNA3, which has the guide RNA scaffold sequence shortened to 100 nt; and gRNA4, which has the guide RNA scaffold sequence shortened to 77 nt. [Figure 12] FIG. 12 shows the effect of point mutations in modified T8-2. [Figure 13] Figure 13 shows the increased production of target protein by an inactivating mutation-containing TnpB-like RNA-guided nuclease complex, which comprises an inactivating mutation-containing T8-2, a TnpB-like RNA-guided nuclease fused with a P300 core domain, and a guide RNA. DETAILED DESCRIPTION OF THE INVENTION

[0011] One aspect of the present invention relates to a TnpB-like RNA-guided nuclease complex, a kit for expressing the complex, a method for cleaving target DNA using the TnpB-like RNA-guided nuclease complex, and a genome editing method.

[0012] [TnpB-like RNA-guided nuclease complex] The TnpB-like RNA-guided nuclease complex of the present invention comprises: (1) TnpB-like RNA-guided nuclease, and (2) gRNA consisting of a sequence including a guide sequence and a gRNA scaffold sequence Such a complex may cleave a specific DNA in a test tube, or may be formed in a cell to cleave a specific site in DNA complementary to the guide RNA sequence. The complex can be formed by genetically introducing (1) a TnpB-like RNA-guided nuclease and (2) a vector containing an expression cassette encoding the guide RNA, messenger RNA, or the guide RNA itself into a cell and expressing it. Alternatively, a complex previously formed in vitro may be introduced into a cell using techniques such as liposomes or electroporation.

[0013] [TnpB-like RNA-guided nuclease] The TnpB-like RNA-guided nuclease (hereinafter also referred to as T protein) that forms the TnpB-like RNA-guided nuclease complex comprises one or more of the following characteristics: (i) The following regular expression array: YKGRTF-[NS]-[KR]-[LM]-x-[AN]-xG-[AS]-[KR]-[GS]-QYx(2)-R-[AS]-x-[DN]-xLxWxG (SEQ ID NO: 1) (In the formula, [NS] represents N or S; [KR] represents K or R; [LM] represents L or S, x represents any amino acid, [AN] represents A or N, [AS] represents A or S, [KR] represents K or R; [GS] represents G or S, x(2) represents any two amino acids; [AS] represents A or S, [DN] stands for D or N) Includes; (ii) comprising at least one domain selected from the group consisting of RuvCI, bridge helix, RuvCII, RuvCIII, and Wedge; (iii) has a length of 300 to 700 residues, preferably 500 to 675 residues, and more preferably 550 to 650; and (iv) It has RNA-guided nuclease activity.

[0014] The regular expression sequence in the present invention is a conserved sequence among TnpB-like RNA-guided nucleases of clades 1 to 4. The regular expression sequence is a 29-amino acid sequence contained in the RuvC domain (more specifically, RuvCIII) of TnpB-like RNA-guided nucleases. By performing a database search using this regular expression sequence, it is possible to specifically and completely identify T proteins.

[0015] More specifically, the TnpB-like RNA-guided nuclease has the following regular expression sequence: (i) the following: YKGRTFNKMINNGSKGQYNxR-SxNxLKWRG (SEQ ID NO: 109) (wherein x independently represents any amino acid.) and is 300 to 700 amino acids in length. Such a regular expression sequence refers to a TnpB-like RNA-guided nuclease that is included in Clade 1 and does not include T11, T16, or T18. Specifically, it may refer to TnpB-like RNA-guided nucleases designated as T6, T8-1 to T8-4, T15, T17, and T21-1 to T21-4, for example.

[0016] The bridge helix, RuvCI, RuvCII, and RuvCIII are each thought to be involved in the endonuclease domain, and Cas, IscB, and TnpB also have similar domains. Meanwhile, the regular expression sequence in RuvCIII is shared only by TnpB-like RNA-guided nucleases and does not match that of Cas9, Cas12, IscB, and TnpB. The TnpB-like RNA-guided nuclease of the present invention may contain all of the domains of the bridge helix, RuvCI, RuvCII, and RuvCIII in order to exert RNA-guided nuclease activity. Furthermore, it may be identified by including a consensus sequence in each domain.

[0017] The RNA-guided nuclease activity of TnpB-like RNA-guided nucleases recognizes a 3- to 7-base sequence, particularly a 4- or 5-base sequence called a TAM (Transposon Associated Motif) sequence, and cleaves the DNA of the target sequence downstream. Therefore, the target sequence to which the guide RNA binds is designed to be adjacent to the TAM sequence, and the TnpB-like RNA-guided nuclease targets and cleaves the target nucleic acid containing the target sequence and the TAM sequence. While TAM sequences may differ for each TnpB-like RNA-guided nuclease, the same TAM sequence can usually be used within the same clade (Figure 5). Furthermore, TAM sequences can be determined by methods well known in the art. For example, a random TAM library in which a random site of several bases, e.g., 7 bases, is linked to the target sequence is treated with a complex of the guide RNA of the target sequence and the TnpB-like RNA-guided endonuclease, and cleavage is confirmed. As an example, typical TAM sequences recognized by TnpB-like RNA-guided nucleases of each clade are as follows: [Table 1]

[0018] TnpB-like RNA-guided nucleases, specifically referred to as T proteins (e.g., T1-T21), can be phylogenetically classified into clades 1 to 5. Phylogenetic classification can be performed using methods well known in the art. For example, homologs of TnpB-like RNA-guided nucleases can be searched for using NCBI's PHI-BLAST with the amino acid sequence regular expression (SEQ ID NO: 1) or the clade 1 regular expression (SEQ ID NO: 109) as a seed query. The results can then be used to search for and collect proteins with sequence homology using BLASTP software, and alignments such as ClustalW can be performed to create a phylogenetic tree from the alignment scores. As another example, the amino acid sequence of T8-2 protein (SEQ ID NO: 22) can be used as a seed query to detect non-redundant protein sequences using NCBI BLASTP, and the phylogenetic tree can be determined based on phylogenetic tree analysis of protein sequences using the neighbor-joining method. As another example, using a partial three-dimensional structure of the RuvC III domain as a query in FoldSeek, proteins can be classified into protein families based on the RMSD score. The present invention particularly relates to T proteins of clades 1 to 4. Proteins included in Clade 1 include T6, T8-1 to T8-4, T11, T15, T16, T17, T18, and T21-1 to T21-4. Among them, Clade 1 is characterized by excluding T11, T16, and T18, and relates to TnpB-like RNA-guided nucleases represented by T6, T8-1 to T8-4, T15, T17, and T21-1 to T21-4. Proteins in clade 2 include T2, T3, and T19. Proteins in clade 3 include T4-1, T4-2, and T5. Proteins in clade 4 include T1, T13, and T14. Proteins in clade 5 include T7, T9, T10, T12, and T20. Phylogenetically, clade 5 is most closely related to known TnpBs, followed by clade 4, clade 3, clade 2, and clade 1 in that order.

[0019] The amino acid sequences and ORF nucleotide sequences of proteins belonging to clade 1 are as follows: [Table 2]

[0020] The amino acid sequences and ORF nucleotide sequences of proteins belonging to clade 2 are as follows: [Table 3]

[0021] The amino acid sequences and ORF nucleotide sequences of proteins belonging to clade 3 are as follows: [Table 4]

[0022] The amino acid sequences and ORF nucleotide sequences of proteins belonging to clade 4 are as follows: [Table 5]

[0023] The amino acid sequences and ORF nucleotide sequences of proteins belonging to clade 5 are as follows: [Table 6]

[0024] In one embodiment of the present invention, the TnpB-like RNA-guided nuclease of the present invention is a protein having an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 49, 50, 52, 54, and 56; or It relates to the mutant protein. A mutant protein refers to a protein that has an amino acid sequence that has at least 50% sequence homology or identity to the amino acid sequence of the original protein and has TnpB-like RNA-guided nuclease activity. Such mutant proteins have the following canonical sequence in the RuvC domain: YKGRTF-[NS]-[KR]-[LM]-x-[AN]-xG-[AS]-[KR]-[GS]-QYx(2)-R-[AS]-x-[DN]-xLxWxG (SEQ ID NO: 1) (In the formula, [NS] represents N or S; [KR] represents K or R; [LM] represents L or S, x represents any amino acid, [AN] represents A or N, [AS] represents A or S, [KR] represents K or R; [GS] represents G or S, x(2) represents any two amino acids; [AS] represents A or S, [DN] stands for D or N) The RuvC domain has a total length of 200 to 350, preferably 250 to 300. Such mutant proteins have steric structural homology, more preferably structural homology in the RuvC domain.

[0025] A variant protein may be specified as having 40% or more, 45% or more, 50% or more, 55% or more, 60% or more, 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 99% or more homology and / or identity to the amino acid sequence of the original protein. In particular, it is particularly preferred that the mutant protein belongs to the same clade as the original protein. When belonging to the same clade, the homology or identity is usually preferably 50% or more, 60% or more, 70% or more, or 80% or more.

[0026] Clade 1 TnpB-like RNA-guided nucleases include: (1) a protein consisting of an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 16, 20, 22, 24, 26, 32, 40, 42, 46, 50, 52, 54, and 56, particularly a protein consisting of an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 16, 20, 22, 24, 26, 40, 50, 52, 54, and 56, or (2) An amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 16, 20, 22, 24, 26, 32, 40, 42, 46, 50, 52, 54, and 56, particularly an amino acid sequence having at least 50%, 60%, 70%, 75%, 80%, or 90% sequence homology or identity to an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 16, 20, 22, 24, 26, 40, 50, 52, 54, and 56, wherein the sequence is a protein having TnpB-like RNA-guided nuclease activity. The sequence identity can be appropriately selected so as not to include T11, T16, and T18.

[0027] Clade 2 TnpB-like RNA-guided nucleases include: (1) a protein consisting of an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 4, 6, 8, and 48; or (2) A protein having TnpB-like RNA-guided nuclease activity, comprising an amino acid sequence having at least 40%, 50%, or 60% sequence homology or identity to an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 4, 6, 8, and 48. Regarding.

[0028] Clade 3 TnpB-like RNA-guided nucleases include: (1) a protein consisting of an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 10, 12, and 14; or (2) A protein having TnpB-like RNA-guided nuclease activity, comprising an amino acid sequence having at least 80% sequence homology or identity to an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 10, 12, and 14. Regarding.

[0029] Clade 4 TnpB-like RNA-guided nucleases include: (1) a protein consisting of an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 2, 36, and 38; or (2) A protein having TnpB-like RNA-guided nuclease activity, comprising an amino acid sequence having at least 60% sequence homology or identity to an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 2, 36, and 38. Regarding.

[0030] In one embodiment, the TnpB-like RNA-guided nuclease of the present invention is selected from the group consisting of: (1) a protein encoded by a nucleic acid sequence selected from the group consisting of the nucleic acid sequences of SEQ ID NOs: 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 51, 53, 55, and 57; or (2) A protein encoded by a nucleic acid sequence having at least 60% sequence identity to a nucleic acid sequence selected from the group consisting of the nucleic acid sequences of SEQ ID NOs: 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 51, 53, 55, and 57, wherein the protein has TnpB-like RNA-guided nuclease activity. Regarding.

[0031] As used herein, "homology" between two amino acid sequences refers to the ratio of identical or similar amino acid residues appearing at corresponding positions when the two amino acid sequences are aligned, and "identity" between two amino acid sequences refers to the ratio of identical amino acid residues appearing at corresponding positions when the two amino acid sequences are aligned. The "homology" and "identity" between two amino acid sequences can be determined, for example, using the BLAST (Basic Local Alignment Search Tool) program (Altschul et al., J. Mol. Biol., (1990), 215(3):403-10).

[0032] The mutant protein is a protein that has structural homology in the RuvC domain of the original protein, and may also have structural homology over the entire length of the original protein. The location of the RuvC domain varies depending on the type of T protein. For example, in the case of the T8-2 protein, R302 to E580 correspond to the RuvC domain. For other T proteins, the RuvC domain corresponding to such positions can be determined. Structural homology can be determined by the RMSD value and / or TM value when aligned with the three-dimensional structure of the original protein. The RMSD value refers to the root-mean-square deviation of atomic positions. For example, when the C root-mean-square deviation (RMSD) between the backbone of the original protein and the backbone of the mutant is 2.5 A or less, preferably 1.5 A or less, and more preferably 1 A or less, it can be said that there is a high probability that the mutant will exhibit an effect equivalent to that of the original protein. The RMSD value can be determined using software well known in the art, such as PyMOL software align. In addition to or instead of the RMSD value, the TM value (template modeling score) can also be used. Proteins having a TM value of 0.9 or more, preferably 0.91 or more, more preferably 0.93 or more can be said to have structural homology. Therefore, variants can be identified by RMSD value in addition to or instead of identifying them by sequence identity or homology. Specifically, when structural alignment is performed using PyMOL software between a T8-2 protein belonging to clade 1 and a TnpB-like RNA-guided nuclease belonging to clades 1 to 4, the RMSD value is 1.9 or less and the TM value is 0.9 or greater. In another example, when structural alignment is performed using PyMOL software between a T8-2 protein belonging to clade 1 and a TnpB-like RNA-guided nuclease belonging to clade 1, the RMSD value is 1.5 or less and the TM value is 0.91 or greater. Since all proteins within a clade have TnpB-like RNA-guided nuclease activity, if the RMSD between the original protein and the mutant protein is 1.9 or less, preferably 1.5 or less, and more preferably 1 or less, and the TM value is 0.9 or greater, preferably 0.91 or greater, and more preferably 0.93 or greater, the mutant protein is presumed to have equivalent activity to the original protein and can be used in genome editing tools using the procedures described herein.

[0033] In another example, the three-dimensional structures and homology of the T8-2 protein belonging to Clade 1 are compared with those of known highly active TnpB RNA-guided nucleases. Known highly active TnpB RNA-guided nucleases include DraTnpB-AI (SEQ ID NO: 174). Aligning the three-dimensional structures allows identification of amino acid positions that contribute to activity. In the amino acid sequence of the T8-2 protein (SEQ ID NO: 22), the amino acid positions corresponding to Q at position 284, R at position 458, K at position 89, S at position 72 and / or S at position 75, R at position 291, S at position 201, L at position 278, and S at position 315 may contribute to the activity of TnpB-like RNA-guided nucleases belonging to Clade 1. The corresponding amino acid positions can be determined by structural or amino acid sequence comparison. In one embodiment, the present invention relates to an amino acid sequence of a TnpB-like RNA-guided nuclease belonging to Clade 1, comprising a substitution at at least one position or a combination thereof selected from the group consisting of Q at position 284, R at position 458, K at position 89, S at position 72 and / or S at position 75, R at position 291, S at position 201, L at position 278, and S at position 315 in the amino acid sequence of SEQ ID NO: 22. These mutations are particularly preferred in combinations of two or three positions. For example, a combination of K at position 89 and Q at position 284, K at position 89 and R at position 291, Q at position 284 and R at position 291, or a combination of K at position 89, Q at position 284, and R at position 291 is preferred.

[0034] Such substitutions include Q284R or Q284K when there is a substitution at position 284 or a corresponding position, K89R when there is a substitution at position 89 or a corresponding position, S72H, S72N, S72Q, S72K, S72R, S75H, S75N, S75Q, S75K, or S75R when there is a substitution at position 72 or 75 or a corresponding position, R291E or R291K when there is a substitution at position 291 or a corresponding position, Q201H or Q201S when there is a substitution at position 201 or a corresponding position, L278K or L278R when there is a substitution at position 278 or a corresponding position, and S315A or S135V when there is a substitution at position 315 or a corresponding position. These substitutions can enhance TnpB-like RNA-guided nuclease activity. More specifically, from the viewpoint of achieving high genome editing efficiency, substitutions of K89R, Q284K, Q284R, and R291K, or combinations thereof, are preferred. For example, substitutions of K89R and Q284K, K89R and R291K, Q284K and R291K, or K89R, Q284K, and R291K are preferred (Figure 12).

[0035] The TnpB-like RNA-guided nuclease or its mutant protein of the present invention has TnpB-like RNA-guided nuclease activity over a wide temperature range, for example, 20 to 50°C, but is particularly active at 30 to 45°C, and more preferably around 37°C (Figure 8). Typical TnpB proteins and TnpB-derived proteins often exhibit maximum activity at around 40 to 50°C, allowing for efficient use at lower temperatures. Because they are highly active not only around 37°C, the body temperature of warm-blooded animals, but also at 25°C, they are particularly suitable for genome editing in living organisms such as plants, algae, mollusks, and microorganisms.

[0036] [Guide RNA (gRNA)] In the present invention, guide RNAs are designed based on the DNA sequence of the cleavage target and the TnpB-like RNA-guided nuclease used, and are configured to include a guide sequence and a gRNA scaffold sequence. Because the DNA of the cleavage target is cleaved near the TAM sequence, it is selected under the condition that it contains a TAM sequence. A 14- to 25-nt sequence adjacent to the 3' end of the TAM sequence can be selected as the guide sequence. The guide sequence may match the DNA sequence of the cleavage target, or it may contain one or several mismatch sequences. Even if the guide sequence does not perfectly match the target sequence, cleavage may occur due to off-target effects. The gRNA scaffold must have a sequence that can be used by the TnpB-like RNA-guided nuclease used. Typically, the 3'-end boundary of the locus (i.e., the boundary between the gRNA scaffold and the guide sequence) can be determined by sequence comparison of each TnpB-like RNA-guided nuclease at the locus. The four bases at the 3' end of the gRNA scaffold sequence can be defined as a transposon end motif (TEM). The gRNA scaffold sequence may be 70 to 200 nt. The guide sequence may be designed to be adjacent to the 3' end of the scaffold sequence, i.e., the TEM sequence, or any sequence of 20 to 100 nt may be inserted.

[0037] Within each clade, the gRNA scaffold sequences of TnpB-like RNA-guided nucleases have high homology / identity. For example, within the same clade, gRNA scaffold sequences can be used if they have 70% or more, preferably 80% or more, and more preferably 90% or more identity / identity. Representative gRNA scaffold sequences for each grade are as follows: [Table 7] The underlined parts represent the TEM sequence. Therefore, the gRNA scaffold can be any of the sequences of SEQ ID NOs: 58 to 62, or a sequence having 70% or more, preferably 80% or more, and more preferably 90% or more sequence identity to the sequence.

[0038] Regions of the gRNA scaffold sequence that are more important for activity can be identified based on sequence identity within the clade, and shortened sequences containing such important regions can also be created. As an example, shortened sequences of the T8-2 gRNA scaffold sequence include SEQ ID NO: 117 (gRNA1), SEQ ID NO: 118 (gRNA2), and SEQ ID NO: 119 (gRNA3). Alternatively, shortened mutant sequences that have at least 90%, more preferably at least 92%, at least 93%, 95%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 119 (gRNA3) and can function as RNA scaffolds can also be selected.

[0039] A gRNA designed based on the TnpB-like RNA-guided nuclease used and the nucleic acid sequence of the cleavage target can be expressed using techniques well known in the art. For example, a nucleic acid containing an expression cassette in which DNA containing a gRNA sequence is placed under the control of a promoter can be used. Such an expression cassette may contain a transcription termination sequence in the 3' region. The nucleic acid containing the expression cassette is incorporated into a vector or plasmid and introduced into a specific cell in vitro or to produce the gRNA.

[0040] [TnpB-like RNA-guided nuclease containing inactivating mutations] TnpB-like RNA-guided nucleases can be inactivated by substituting amino acid residues in their activation domain. Examples of activation domains include D308, E489, and D575 in T8-2. Substituting the residues at these positions can inactivate TnpB-like RNA-guided nucleases. Inactivated TnpB-like RNA-guided nucleases are recruited to target sequences by the action of guide RNA, but do not exhibit RNA-guided nuclease activity. By fusing a functional domain to a TnpB-like RNA-guided nuclease containing an activating mutation, its function on genes in target sequences can be regulated. The functional domain to be fused may be any domain; examples include transcriptional regulatory domains such as p300, DMNT3, KRAB, SRDX, VPR, and VP64. In another example, any functional domain, such as APOBEC, AID, TadA, Dda, FokI, M-MLV-RTase, or GFP, can be used, enabling not only epigenetic state but also base modification, reverse transcription, phosphate bond cleavage, and photomanipulation of fluorescent proteins. RNA-guided nucleases containing inactivating mutations fused to functional domains are also widely known, including CRISPR-Cas RNA-guided nucleases (Non-Patent Document 10: Int J Mol Sci., 2020 Apr 25;21(9):3038. doi: 10.3390 / ijms21093038), and similar functional regulation is possible with TnpB-like RNA-guided nucleases.

[0041] A TnpB-like RNA-guided nuclease containing an inactivating mutation fused to a functional domain forms a complex with a guide RNA consisting of a sequence including a guide sequence and a gRNA scaffold sequence, and exerts a function corresponding to the functional domain. When a transcription-promoting domain, such as all or a portion of the functional domain of p300, VPR, or VP64, is fused as the functional domain, transcription in the region recruited by the guide RNA is promoted. On the other hand, when a transcription-repressing domain, such as all or a portion of the functional domain of DMNT3, KRAB, or SRDX, is fused as the functional domain, transcription in the region recruited by the guide RNA is suppressed. Therefore, the present invention may relate to a method for regulating transcription of target DNA using a TnpB-like RNA-guided nuclease containing an inactivating mutation fused to such a functional domain.

[0042] [Target-specific cleavage method] Another aspect of the present invention relates to a method for cleaving a target DNA using the TnpB-like RNA-guided nuclease complex of the present invention. More specifically, the method comprises the steps of: (i) a TnpB-like RNA-guided nuclease according to the present invention; (ii) a guide RNA comprising a sequence including a guide sequence and a gRNA scaffold sequence; The present invention relates to a method for cleaving target DNA using a complex comprising the complex. The complex may be formed in vitro and then introduced into cells using electroporation, liposomes, or the like to cleave the target DNA, or the target DNA may be cleaved in vitro. Alternatively, the complex may be incorporated into an expression cassette and transfected into cells using a vector, messenger RNA, or guide RNA itself, and the TnpB-like RNA-guided nuclease and guide RNA may be expressed in the cells to cleave genomic DNA. Such a cleavage method can be performed using a kit for cleaving target nucleic acids that includes an expression cassette.

[0043] Gene transfer can be carried out using techniques known in the art. The transfer method is not particularly limited and can be selected appropriately depending on the type of material to be transferred and the target of transfer. Transfer methods are broadly divided into direct methods and methods using viral vectors. Direct methods include electroporation, liposomes, particle gun injection with gold particles, and whisker transfer. For methods using viral vectors, adenovirus, adeno-associated virus, lentivirus, Agrobacterium, tobacco mosaic virus (TMV), and the like can be used as vectors depending on the host species.

[0044] Target-specific cleavage occurring within cells can be repaired by genomic DNA repair mechanisms. Non-homologous end joining or homologous recombination occurs during this process, enabling gene knockout or knock-in. Therefore, the target-specific cleavage method of the present invention can also be referred to as a genome editing method. The genome editing method of the present invention may be performed on cultured cells, cultured tissues, or living cells. The target species is not particularly limited, but genome editing is possible for any bacterial, archaeal, or eukaryotic cell. Eukaryotic cells may include any cell, such as plant cells, insect cells, or animal cells. For example, genome editing is possible for any mammalian cell, and genome editing is possible for both human and non-human animal cells. The T protein of the present invention is characterized by a low optimal temperature. Therefore, genome editing can be performed at temperatures of 10 to 30°C.

[0045] [Target-specific protein recruitment] Another aspect of the present invention relates to a method for recruiting a specific protein to a target DNA using the TnpB-like RNA-guided nuclease complex of the present invention, which is engineered so that it does not exhibit TnpB-like RNA-guided nuclease activity. As a complex engineered to not exhibit TnpB-like RNA-guided nuclease activity, the activity of the TnpB-like RNA-guided nuclease may be inactivated, or the length of the guide RNA may be adjusted to prevent the TnpB-like protein from exhibiting DNA cleavage activity while containing a sequence long enough to specifically bind to the target. Such a guide RNA may be a guide sequence of a target sequence of about 10 bases. More specifically, the TnpB-like RNA-guided nuclease complex used in the method for recruiting a specific protein to target DNA may be as follows: (i) a TnpB-like RNA-guided nuclease according to the present invention in which the amino acid residue at the active center is substituted with another amino acid; (ii) a guide RNA comprising a sequence including a guide sequence and a gRNA scaffold sequence; (iii) specific proteins that bind to or interact with TnpB-like RNA-guided nucleases; By expressing or introducing such a complex into cells, the complex is recruited to the vicinity of the guide sequence site, but cleavage does not occur, and specific proteins bound to the complex can be recruited. The active center of the mutant TnpB-like RNA-guided nuclease described in (i) can be a typical DED active center contained in the RuvC domain. In the case of T8-2, the activity of the TnpB-like guided nuclease can be inactivated by introducing mutations into D308, E489, and D575. The mutation can be selected appropriately within the range that inactivates the activity, and an example is a mutation to alanine. The RNA scaffold sequence of the guide RNA described in (ii) can be inserted with binding sites for other proteins, such as MS2 sequences. The specific proteins described in (iii) can modify DNA or epigenetic states, specifically, any protein commonly referred to as an epigenetic factor. More specifically, p300, DMNT3, KRAB, SRDX, VPR, VP64, etc. can be used. In addition to the epigenetic state, proteins that modify bases, reverse transcription, cleave phosphate bonds, and bind light-manipulation tools such as fluorescent proteins (specifically, APOBEC, AID, TadA, Dda, FokI, M-MLV-RTase, GFP, etc.) can also be used. By recruiting such specific proteins to the target DNA region, the epigenetic state of the target DNA region can be altered.

[0046] [kit] Another aspect of the present invention may relate to a kit for cleaving a target nucleic acid in a cell or a kit for genome editing. (i) an expression cassette for the TnpB-like RNA-guided nuclease of the present invention; (ii) an expression cassette for a guide RNA having a sequence including a gRNA scaffold sequence; The guide RNA expression cassette is provided so that a guide sequence can be inserted. A guide RNA expression cassette can be prepared by designing a cleavage target sequence and introducing the guide sequence. The expression cassette for the TnpB-like RNA-guided nuclease and the expression cassette for the guide RNA with the introduced guide sequence are incorporated into, for example, a plasmid or vector, and then genetically introduced into cells.

[0047] Expression cassettes typically contain a promoter sequence and can be prepared by placing a sequence encoding a TnpB-like RNA-guided nuclease or guide RNA under its control. The promoter can transiently or constitutively control the expression of downstream sequences. The expression cassettes may be contained on a single polynucleotide or on different polynucleotides. The expression cassette may also contain elements contributing to expression, such as a terminator, as well as elements necessary for the preparation of a plasmid or vector, such as a signal sequence, a tag sequence, a reporter sequence, a multicloning site, a drug resistance gene, and an origin of replication, as long as the activity of the TnpB-like RNA-guided nuclease is not impaired. To ensure that the expressed protein functions within the nucleus, it is desirable to add a nuclear localization signal to the signal sequence. One or more nuclear localization signals may be arranged in series. The nuclear localization signal can be selected appropriately depending on the organism. This allows the TnpB-like RNA-guided nuclease expressed in the cell to translocate into the nucleus, where it can cooperate with the guide RNA to cleave the target nucleic acid sequence.

[0048] The genome editing kit of the present invention may further include other materials, reagents, tools, etc., necessary for carrying out the genome editing method of the present invention, such as a nucleic acid introduction reagent and a buffer solution, as needed. Other materials necessary for carrying out the genome editing method of the present invention include a donor polynucleotide in addition to an expression cassette for a TnpB-like RNA-guided nuclease and / or an expression cassette for a guide RNA. Introducing the donor polynucleotide into the nucleus enables knock-in of the donor polynucleotide at the cleavage site created by the CRISPR / Cas system. The donor polynucleotide causes homologous recombination at the cleavage site by locating sequences homologous to the 5' and 3' sequences of the introduced sequence, respectively.

[0049] All documents mentioned herein are incorporated by reference in their entirety.

[0050] The present invention will be described in more detail below with reference to examples. However, these examples are merely examples shown for the convenience of explanation, and the present invention is not limited to these examples in any sense. [Example]

[0051] Example 1: Identification of TnpB-like RNA-guided nuclease homologs TnpB-like RNA-guided nuclease homologs were detected using NCBI BLASTP and TBLASTN against NCBI's non-redundant protein sequences, using the amino acid sequence of the T8-2 protein (SEQ ID NO: 22) as a seed query. Similar searches were also performed in the MGNIFY database as well as NCBI. The nucleotide sequences of the loci corresponding to each TnpB-like RNA-guided nuclease were obtained from the NCBI GenBank database. Based on phylogenetic analysis of protein sequences using the neighbor-joining method, TnpB-like RNA-guided nucleases were classified into five groups (clades 1 to 5) (Figure 1). The rightmost boundary of the locus (i.e., the boundary between the guide RNA (gRNA) scaffold and the guide sequence) was determined by aligning the genomic sequences of the loci encoding the T proteins and comparing their 3' ends using ClustalW (https: / / www.genome.jp / tools-bin / clustalw). From the alignment, the guide-scaffold boundary was identified as the most downstream position where sequence conservation dropped off sharply (Fig. 2). A sequence motif consisting of four nucleotides located at the 3' end of the scaffold was defined as a transposon-encoded motif (TEM).

[0052] By aligning the amino acid sequences of TnpB-like RNA-guided nucleases in clades 1 to 4, we determined the conserved sequences (canonical expressions) among these sequences: YKGRTF-[NS]-[KR]-[LM]-x-[AN]-xG-[AS]-[KR]-[GS]-QYx(2)-R-[AS]-x-[DN]-xLxWxG (SEQ ID NO: 1) (In the formula, [NS] represents N or S; [KR] represents K or R; [LM] represents L or S, x represents any amino acid, [AN] represents A or N, [AS] represents A or S, [KR] represents K or R; [GS] represents G or S, x(2) represents any two amino acids; [AS] represents A or S, [DN] stands for D or N) By performing a database search using PHI-BLAST or the like on this regular expression sequence, it is possible to specifically and completely identify clades 1-4 of T proteins.

[0053] Example 2: Construction of TnpB-like RNA-guided nuclease expression vector To produce recombinant TnpB-like RNA-guided nucleases in E. coli as N-terminal 10x histidine-tagged maltose-binding protein-TEV protease cleavage site (His-MBP-TEV) fusion proteins, a synthetic DNA fragment containing the His-MBP sequence was cloned between the NcoI and NdeI cleavage sites of pET28b using the NEBuilder HiFi DNA Assembly Kit (New England Biolabs). This resulted in the pAN36 vector. All genes encoding TnpB-like RNA-guided nucleases and the associated gRNA scaffolds were synthesized by Integrated DNA Technology (IDT) (Table 8). For some synthetic TnpB-like RNA-guided nuclease genes, the 5' region of the coding sequence that did not overlap with the gRNA scaffold was codon-optimized for protein expression in E. coli. This DNA fragment was cloned into the pAN36 vector under the control of a T7 promoter. In the resulting vector, the sequence encoding the TnpB-like RNA-guided nuclease was fused in-frame to the N-terminal His-MBP-TEV sequence. A description of the His-MBP-TEV-TnpB-like RNA-guided nuclease expression vector is shown in Table 8. [Table 8]

[0054] Example 3: Expression and purification of TnpB-like RNA-guided nuclease-RNA complexes in E. coli For expression of the recombinant His-MBP-TEV fusion TnpB-like RNA-guided nuclease, Escherichia coli Rosetta2(DE3)pLysS strain (Novagen) was transformed with the His-MBP-TEV-fusion TnpB-like RNA-guided nuclease expression vector. E. coli was cultured overnight at 37°C in Luria-Bertani (LB) medium supplemented with 50 μg / ml kanamycin and 34 μg / ml chloramphenicol. 10 ml of the overnight culture was inoculated into 1 liter of LB medium supplemented with 50 μg / ml kanamycin and 34 μg / ml chloramphenicol, and the optical density (OD) was measured.600 The cells were grown until the chromatin density (DDS) reached 0.6. Gene expression was then induced by adding 0.25 mM IPTG and grown at 18°C ​​for 20 hours. The cells were harvested by centrifugation and stored at -70°C until use.

[0055] All subsequent purification steps were performed at 4°C. Cells were resuspended in buffer A (50 mM Tris-HCl, pH 8.0, 500 mM NaCl, 5% (v / v) glycerol, and 25 mM imidazole) supplemented with bovine DNAse I (Fujifilm Wako Pure Chemical Industries, Ltd.), chicken lysozyme (Fujifilm Wako Pure Chemical Industries, Ltd.), and protease inhibitors (phenylmethylsulfonyl fluoride and Roche Complete ethylenediaminetetraacetic acid-free). After a 30-minute incubation, cells were disrupted by sonication (ULTRASONIC DISRUPTOR UD-211, TOMY) on ice. After centrifugation at 40,000 g for 30 min to remove cellular debris, the supernatant was filtered through a 0.45 μm polyvinylidene fluoride (PVDF) membrane and batch-coupled to 1 ml of Ni-Sepharose 6 Fast Flow Resin (Cytiva) equilibrated with buffer A for 1 h. The resin was loaded onto an Econo-Pac chromatography column (Bio-Rad) and washed first with 14 ml of buffer B (50 mM Tris-HCl, pH 8.0, 1 M NaCl, 5% (v / v) glycerol, 25 mM imidazole) and then with 10 ml of buffer A. The bound protein was eluted with 3.5 ml of buffer C (50 mM Tris-HCl, pH 8.0, 1 M NaCl, 5% (v / v) glycerol, 300 mM imidazole). The proteins in the eluate were separated by 5-20% (w / v) polyacrylamide gel SDS-PAGE, and the gel was stained with Coomassie Brilliant Blue (Figure 3: SDS-PAGE gel). The peak fractions containing the fusion protein were transferred to nuclease-free tubes, flash-frozen in liquid nitrogen, and stored at -70 °C until use. The resulting TnpB-like RNA-guided nuclease ribonucleoprotein (RNP) samples were used for nucleic acid extraction and dsDNA cleavage analysis. For in vitro double-stranded DNS cleavage analysis, TnpB-like RNA-guided nuclease RNP samples were further purified on a HiLoad 16 / 600 Superdex 200 pg column in buffer D [25 mM Tris-HCl pH 8.0, 150 mM NaCl, 5 mM MgCl2, 1% (v / v) glycerol, 1 mM DTT]. Pooled fractions were concentrated by ultrafiltration and stored at -80°C until further use.

[0056] Example 4: Extraction and analysis of gRNA bound to TnpB-like RNA-guided nucleases To extract RNA bound to TnpB-like RNA-guided nuclease RNPs, 180 μl of the peak fraction containing RNPs was vigorously mixed with 540 μl of TRI Regent (Molecular Research Center, Inc.) and 108 μl of chloroform for 15 seconds, incubated at room temperature for 5 minutes, and then centrifuged at 12,000 g for 15 minutes at 4°C. The upper aqueous phase containing the RNA extracted from the RNPs was transferred to a new tube and mixed with 500 μl of 100% (v / v) ethanol. The sample was then loaded onto an RNA Clean & Concentrator-5 spin column (Zymo Research). After washing the spin column with wash buffer, 13.6 units of DNAse (QIAGEN) were applied to the spin column and incubated at room temperature for 15 minutes to remove residual DNA. After washing the spin column three times with buffer, the bound RNA was eluted with 15 μl of nuclease-free water. 500 ng of purified RNA was then used for RNA library preparation. RNA libraries were prepared using the SMARTer smRNA-Seq Kit for Illumina (TAKARA-BIO) according to the manufacturer's instructions. The resulting NGS adapter-ligated cDNA was amplified by eight cycles of PCR using full-length Illumina NGS indexing primers. Amplified DNA fragments of 200–500 bp were size-selected by 3% (w / v) agarose gel electrophoresis and gel-extracted and purified using the FastGene Gel / PCR Extraction Kit (FastGene). DNA fragments were pooled in equimolar ratios and adjusted to a 50 pM concentration using the Qubit dsDNA HS Assay Kit and Qubit 3.0 Fluorometer (Thermo Fisher Scientific Inc.). The pooled library was mixed with 0.25 volumes of 50 pM PhiX control v3 (Illumina Inc.) and then used for 2 × 150 bp paired-end sequencing on an iSeq100 (Illumina Inc.). Reads were adapter trimmed and aligned to the template sequence using Bowtie2 software. The identified sequences of small RNAs bound to each TnpB-like RNA-guided nuclease were defined as the sequences of the corresponding gRNAs (Figure 4, a diagram of CDS and NGS reads fused together). Information about the gRNAs is summarized in the table below. [Table 9]

[0057] Example 5: Construction of 7N-TAM library plasmids The TAM sequence of TnpB-like RNA-guided nucleases was determined using a plasmid library containing seven randomized nucleotides (7N). To generate the 7N-TAM plasmid library, ssDNA containing seven randomized nucleotides (5'-AGCTATGACCATGATTACGAATTCNNNNNNNCTGCAGGAGCAAAGACC-3' (SEQ ID NO: 87)) was converted to dsDNA using reverse transcriptase (ReverTra Ace, Takara) with the primer (5'-GTCTTTGCTCCTGCAG-3' (SEQ ID NO: 88)). This dsDNA fragment was assembled with pUC18 containing an NGS adapter sequence using the NEBuilder HiFi DNA Assembly Cloning Kit (NEB) to generate the 7N-TAM plasmid library (Table 10, plasmid sequence). E. coli DH5α cells were transformed with the reaction mixture and grown at 37°C on LB plates supplemented with 100 μg / ml ampicillin. More than 100,000 colonies were washed from the plate, and a plasmid library was extracted using a NucleoBond Xtra Midiprep kit (Takara). [Table 10]

[0058] Example 6: Screening of TAM sequences by in vitro transcription / translation In vitro transcription / translation (IVTT) reactions were performed using the PUREfrex 2.0 Reconstituted Cell-Free Protein Synthesis Kit (GeneFrontier). According to the manufacturer's instructions, a DNA template encoding a T7 promoter-driven TnpB-like RNA-guided nuclease and a T7 promoter-driven gRNA scaffold with a spacer sequence (20 nt: 5'-GAATTCGTAATCATGGTCAT-3' (SEQ ID NO: 90)) targeting the TAM library was generated by PCR using the synthetic gene encoding the TnpB-like RNA-guided nuclease as a template. The PCR primers used for DNA template preparation were Illumina's Nextera HT v2 dual index primers. IVTT reactions were performed in 20-25 μl reactions using 100 ng of TnpB-like RNA-guided nuclease template, 125 ng of the corresponding gRNA template with the spacer sequence, and 60 ng of the TAM library plasmid. An IVTT reaction without the gRNA DNA template was used as a control. The reaction was incubated at 37°C for 4 hours, then quenched by adding 5 μl of 100 mg / ml RNase A (Nippon Gene) and incubated at 37°C for 30 minutes. After adding 5 μl of 1% SDS, the reaction was incubated at 50°C for 10 minutes to denature the protein. Four units of protease K (FUJIFILM) were added and incubated at 37°C for 90 minutes to digest the protein. The undigested TAM library plasmid was recovered from the reaction using the FastGene Gel / PCR Extraction Kit and used as a template for subsequent PCR. A 250-bp DNA fragment containing the TAM motif and target sequence was PCR-amplified for 30 cycles using KOD One PCR Master Mix (Toyobo) and Illumina full-length Nextera NGS indexing primers. The amplified DNA fragment was separated by 2.5% (w / v) agarose gel electrophoresis and then gel-extracted and purified. Preferentially depleted TAM motifs were identified by NGS amplicon sequencing of the TAM library plasmid as described above. TAM preferences were characterized from FASTQ files using a custom Python script.Briefly, 7-nucleotide TAMs were extracted and counted. TAM frequencies were normalized to the sequencing depth of each sample. Sequence displays were generated using WebLogo version 2.8.2 (http: / / weblogo.berkeley.edu / ) using TAMs with at least a 10-fold decline relative to the control (Figure 5).

[0059] Example 7: In vitro dsDNA cleavage analysis The 7N sequence was converted to a TAM sequence by site-directed mutagenesis, and substrate plasmids containing the TAM sequence were synthesized. Substrate plasmids containing different target sequences (5'-GTTCTCCAGGCTGCTATCCTTAGCA-3' (SEQ ID NO: 91) and 5'-GAATTCGTAATCATGGTCATAGCTG-3' (SEQ ID NO: 92)) were also synthesized by site-directed mutagenesis. A 716-bp substrate dsDNA was amplified by PCR using the TAM sequence plasmid and primer 1 (5'-AAAGGGGATGTGCTGCAAGG-3' (SEQ ID NO: 93)) and primer 2 (5'-TATCTTTATAGTCCTGTCGG-3' (SEQ ID NO: 94)). The amplified dsDNA fragment (20 nM) was mixed with purified TnpB-like RNA-guided nuclease RNP (1, 2, or 4 μM) corresponding to the target sequence in 5 μl of buffer D and incubated for 2 hours. The reaction mixture was quenched by adding 2 μl of Proteinase K (Fujifilm Wako Pure Chemical Industries, Ltd.) and incubated at 60 °C for 5 min. The reaction products were analyzed using a MultiNA microchip electrophoresis system (SHIMADZU Inc.) (Figure 6).

[0060] Example 8: In vitro plasmid cleavage analysis The 7N sequence was converted to a TAM sequence by site-directed mutagenesis to synthesize substrate plasmids containing the TAM sequence. Specifically, for T5, 5'-TTAT-3' or 5'-AACAT-3' was added to the 5'-terminus of the target sequence 5'-GAATTCGTAATCATGGTCATAGCTG-3' (SEQ ID NO: 92), and for T8-2, 5'-TTCAT-3' or 5'-AACAT-3' was added to the 5'-terminus of the target sequence. 50 ng of the substrate plasmid was mixed with purified TnpB-like RNA-guided nuclease RNP (4 μM) corresponding to the target sequence in 5 μl of buffer D and incubated at 37°C for 2 hours. The reaction mixture was quenched by adding 2 μl of Proteinase K (Fujifilm Wako Pure Chemical Industries) and incubated at 60°C for 5 minutes. The reaction products were analyzed on a 1.0% (w / v) agarose gel (Figure 7). When the TAM sequence corresponded to the correct one, a change in mobility due to cleavage was observed.

[0061] Example 9: In vitro plasmid cleavage analysis (confirmation of temperature dependency) 50 ng of the substrate plasmid used in Example 8 was mixed with purified TnpB-like RNA-guided nuclease RNP (4 μM) corresponding to the target sequence in 5 μl of buffer D. The reaction temperature was varied from 20°C to 50°C in 5°C intervals, and the mixture was incubated for 2 hours under each condition. The reaction mixture was quenched by adding 2 μl of Proteinase K (Fujifilm Wako Pure Chemical Industries, Ltd.) and then incubated at 60°C for 5 minutes. The reaction products were analyzed on a 1.0% (w / v) agarose gel (Figure 8). These results indicate that the TnpB RNA-guided nuclease used in this invention has a lower active temperature (around 37°C) compared to typical TnpB proteins and TnpB-derived genes, which reach their maximum activity around 40–50°C.

[0062] Example 10: Construction of T8 protein and gRNA expression vectors Plasmid vectors for expressing the T8-1 protein or T8-2 protein and gRNA in cultured human cells were constructed as follows. The plasmid vector for expressing the T protein is a pUC-based transient expression plasmid vector (SEQ ID NO: 95) equipped with a CAG promoter, a multiple cloning site (MCS) that conferred SV40 NLS and Nucleoplasmin NLS to the C-terminus of the T protein, and a bGH polyA signal. The human codon-optimized T8-1 CDS sequence (SEQ ID NO: 96, synthesized by IDT) or the human codon-optimized T8-2 CDS sequence (SEQ ID NO: 97, synthesized by IDT) of the T protein was subcloned into the MCS. The plasmid vector for expressing gRNA is a pUC-based transient expression plasmid vector (SEQ ID NO: 98) that contains a human U6 promoter, a non-coding RNA cloning site, and a hammerhead ribozyme sequence on the 3' side. First, the gT8-1 gRNA scaffold sequence (SEQ ID NO: 99, synthesized by IDT) or the T8-2 gRNA scaffold sequence (SEQ ID NO: 100, synthesized by IDT) was introduced into the gRNA expression vector. Six sequences targeting the hAAVS1 gene (SEQ ID NOs: 101 to 106) were then subcloned.

[0063] Example 11: Genome editing efficiency evaluation using HEK293FT A solution containing 100 ng of the two types of T8 protein expression vectors (expressing T8-1 or T8-2) constructed in Example 10 and 50 ng of gRNA expression vectors (one of each of six types of gRNA expression vectors targeting the hAAVS1 gene) was mixed with Lipofectamine 3000 reagent and transfected into 96-well cultured HEK293FT cells (~10 4 The GFP protein expression plasmid (pEGFP-N1, Clonetech) was added to the control wells of the 12-well transfection wells. After culturing for 3 days, the DNA of the HEK293FT cells was extracted with alkaline buffer (0.1N NaOH), and 100 pg / μL of the roughly purified DNA was used as a template. Using KOD One PCR master Mix and the first round PCR primer set (sequence numbers 110 and 111 when gRNA1, gRNA2, or gRNA3 was used, and sequence numbers 112 and 113 when gRNA4, gRNA5, or gRNA6 was used), a partial sequence of the hAAVS1 gene containing the genome editing target sequence was subjected to a PCR reaction under the following conditions to amplify the target sequence. Similarly, a 20-fold dilution of the first-round PCR solution was used as a template, and a PCR reaction was carried out under the conditions described below using a second-round PCR primer set and second-round PCR primers containing Combinatorial Dual Index (CDI) sequences (5'-AATGATACGGCGACCACCGAGATCTACACNNNNNNNNTCGTCGGCAGCGTC-3' (SEQ ID NO: 107) and 5'-CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTCTCGTGGGCTCGG-3' (SEQ ID NO: 108)). PCR amplification products with adapter and index sequences attached to both ends for sequencing were obtained using an Illumina sequencing device. PCR conditions: 98°C / 2 minutes, 98°C / 10 seconds, 55°C / 5 seconds, 68°C / 3 seconds, 35 cycles (Same for 1st and 2nd rounds) The resulting PCR products were subjected to agarose gel electrophoresis, and the target sequence was excised and purified. The purified DNA was measured for concentration using the Qubit dsDNA HS Assay Kit and a Qubit 3.0 Fluorometer (Thermo Fisher Scientific, Waltham, MA, USA) and adjusted to 50 pM. The pooled library was mixed with 0.25 volumes of 50 pM PhiX control v3 (Illumina) and subjected to 2 × 150-bp paired-end sequencing on an iSeq100 (Illumina, San Diego, CA, USA). The resulting Fastq files were analyzed using CRISPResso2 (Clement et al., 2019 Nature Biotechnology). Since insertions and deletions within 10 bases downstream of the target sequence are considered genome editing induced by T8 protein when the T8 protein exhibits a cleavage pattern similar to that of TnpB, the mutation quantification window width was set to 10 bp from the predicted cleavage site (Figures 9 and 10). The "genome editing efficiency" was calculated as the percentage of all sequenced reads that contained insertions or deletions within the quantitative window width. Figure 9 shows the results of genome editing efficiency evaluation using the above procedure for each of gRNA1-gRNA6, co-expressed with T8-1 or T8-2. The vertical axis shows genome editing efficiency (the percentage of all reads with insertion or deletion mutations in the target region), and the horizontal axis shows the combination of vectors used for transformation. While genome editing efficiency was nearly zero in the control, co-expression with gRNA3 in both T8-1 and T8-2 showed the highest genome editing efficiency (approximately 0.006). gRNA5 was next most efficient (approximately 0.0015). While T8-1 and T8-2 showed similar genome editing efficiencies, gRNA1, gRNA2, and gRNA4 showed higher genome editing efficiencies when co-transfected with T8-2 compared to when co-transfected with T8-1. Figure 10 shows a specific example of genome editing observed when a T8-2 protein expression vector and a gRNA3 expression vector were simultaneously introduced into HEK293FT cells and the genome sequence was analyzed. The target genome sequence is indicated by "Reference," the expected gRNA binding site is indicated by "sgRNA" (rectangle), and the expected cleavage site is indicated by a dashed line. Different sequences output by the sequencer were aligned, and the number of reads and percentage of total reads for each sequence are shown. Base substitutions are indicated in bold, and deletions are indicated by "-." When genome editing occurs, insertions or deletions are observed near the expected cleavage site, indicated by the dashed line. Multiple types of such reads (particularly deletions of 3 to 8 bases) were detected at a rate of approximately 0.1%. These findings demonstrate that HEK293FT cells co-transfected with a T8-2 expression vector and a gRNA3 expression vector undergo genome editing, primarily deletions, in the target genome region.

[0064] Example 12: Construction of modified T8-2 guide RNA expression vector Plasmid vectors for expressing T8-2 guide RNA in human cultured cells were constructed as follows. The plasmid vector for expressing gRNA is a pUC-based transient expression plasmid vector (SEQ ID NO: 98) that contains a human U6 promoter, a non-coding RNA cloning site, and a hammerhead ribozyme sequence on its 3' end. First, a 152-nt T8-2 guide RNA scaffold sequence (SEQ ID NO: 100, synthesized by IDT) was introduced into the gRNA expression vector. Furthermore, a sequence targeting the hAAVS1 gene (SEQ ID NO: 103) was subcloned. This guide RNA, consisting of the T8-2 guide RNA scaffold sequence, target sequence, and hammerhead ribozyme sequence, was designated gRNA original (SEQ ID NO: 114).

[0065] The modified T8-2 guide RNA expression vector was created as follows. A guide RNA expression vector (SEQ ID NO: 115) was created by removing only the hammerhead ribozyme sequence from the transient expression plasmid vector (SEQ ID NO: 98). A guide RNA sequence consisting only of the T8-2 guide RNA scaffold sequence (SEQ ID NO: 98) and the hAAVS1 gene target sequence (SEQ ID NO: 100) was cloned into the cloning site of this vector. This guide RNA sequence was designated ddHDV (SEQ ID NO: 116). Next, nucleotide sequences were gradually deleted from the 5' side of the guide RNA scaffold sequence within the ddHDV sequence, and four types of guide RNAs (synthesized by IDT) with shortened sequences were cloned: gRNA1 (guide RNA scaffold sequence length 131 nt, sequence number 117), gRNA2 (guide RNA scaffold sequence length 111 nt, sequence number 118), gRNA3 (RNA scaffold sequence length 100 nt, sequence number 119), and gRNA4 (guide RNA scaffold sequence length 77 nt, sequence number 120), to create a total of five modified T8-2 guide RNA expression vectors.

[0066] Example 13: Evaluation of genome editing efficiency using modified T8-2 guide RNA using HEK293FT Using the T8-2 protein expression vector constructed in Example 10 and the modified T8-2 guide RNA expression vector constructed in Example 12, the effect of modifying the guide RNA scaffold sequence on genome editing efficiency was evaluated using the method of Example 11 (Figure 11). For comparison, a condition in which a GFP protein expression plasmid (pEGFP-N1, Clonetech) was introduced was prepared (referred to as Control). We compared the genome editing efficiency of the hAAVS1 gene in experiments in which a T8-2 protein expression vector was co-transfected with a gRNA expression vector that shared the same target sequence but had a different gRNA coding region. No insertions or deletions of bases were observed in the hAAVS1 gene in the control vector. When T8-2 was co-expressed with the original gRNA, which contained the 152-nt T8-2 guide RNA scaffold sequence, or with ddHDV, which lacked the HDV sequence, the genome editing efficiency was approximately 0.7%. This result indicated that removing the hammerhead ribozyme sequence from the guide RNA did not affect genome editing efficiency. The genome editing efficiency using gRNA1, which shortened the T8-2 guide RNA scaffold sequence to 131 nt, was also approximately 0.7%. However, using gRNA2, which shortened the guide RNA scaffold sequence to 111 nt, the genome editing efficiency increased to approximately 1.1%, and using gRNA3, which shortened the guide RNA scaffold sequence to 100 nt, the genome editing efficiency improved to approximately 0.86%. However, genome editing activity was not confirmed when gRNA4 was shortened to 77 nt. These results indicate that deleting the 21 nt 5' end of the original T8-2 guide RNA scaffold sequence improves genome editing activity, and furthermore, that the RNA sequences contained in the gRNA scaffolds of gRNA3 and gRNA4 play an important role in maintaining the DNA cleavage activity of T8-2.

[0067] Example 14: Construction of T8-2 mutant protein expression vector As in Example 11, plasmid vectors expressing 24 types of human codon-optimized T8-2 mutant proteins (SEQ ID NOs: 121-168) instead of the T8-2 protein were constructed.

[0068] Example 15: Evaluation of genome editing efficiency by T8-2 mutant protein using HEK293FT The genome editing efficiency of the T8-2 mutant protein was evaluated in the same manner as in Example 12. The plasmid vector used for expressing the gRNA was the plasmid vector containing SEQ ID NO: 118, which had the highest editing efficiency among the six gRNA vectors used in Example 13 (Figure 12). Figure 12 shows the results of genome editing evaluation using the above procedure when unmutated T8-2 protein or each of 24 T8-2 mutant proteins was co-expressed with a gRNA3 expression vector. The vertical axis indicates genome editing efficiency (the percentage of all reads with insertion or deletion mutations in the target region), and the horizontal axis indicates the transformation conditions. Genome editing efficiency was nearly zero for the control and approximately 0.3% for unmutated T8-2 protein. However, when the T8-2 mutant protein contained the K89R, Q284K, or R291K mutations, genome editing efficiency was approximately 1%. Furthermore, genome editing efficiency of approximately 2% was observed for T8-2 mutant proteins containing both the K89R and Q284K mutations, or the K89R and R291K mutations, or the R291K and Q284K mutations. Furthermore, a T8-2 mutant protein containing the K89R, R291K, and Q284K mutations simultaneously showed a genome editing efficiency of approximately 2.5%. Therefore, it was concluded that the K89R, R291K, and Q284K mutations each increase the genome editing efficiency of the T8-2 protein, and that a combination of these mutations can produce a T8-2 mutant protein with improved genome editing efficiency.

[0069] Example 16: Construction of a fusion protein expression plasmid and a gRNA expression vector consisting of inactive T8-2 and the catalytic domain of an epigenetic factor histone acetyltransferase The known TnpB has a typical DED active center sequence contained in the endonuclease domain RuvC. Substitution of the N-terminal aspartic acid with alanine is known to abolish DNA cleavage activity. Furthermore, expression of a fusion protein consisting of a known inactive Cas9 and the catalytic domain of a histone acetyltransferase (P300 core) and a guide RNA targeting a promoter sequence promotes acetylation of histone proteins near the target sequence, thereby inducing transcriptional activation of nearby genes through changes in the epigenetic state of the promoter region. Aiming to utilize T8-2 in gene transcription activation technology, we aimed to create an artificial transcriptional activator consisting of an inactive T8-2 gene created by amino acid substitution in the DED active center sequence and the P300 core. A plasmid vector for expressing a fusion protein consisting of inactive T8-2 and the catalytic domain P300core of histone acetyltransferase and gRNA in human cultured cells was constructed as follows. A human codon-optimized CDS sequence (SEQ ID NO: 170, synthesized by IDT) was created encoding a fusion protein (dead T8-2-P300 core, SEQ ID NO: 169) consisting of an inactive T8-2 sequence in which the 308th aspartic acid residue, which constitutes the DED active center sequence of T8-2, was substituted with alanine, a hemagglutinin tag was attached to the C-terminus of the inactive T8-2, two SV40 NLSs, and the P300 core sequence of human histone acetyltransferase. The fusion protein expression vector was constructed by cloning the CDS sequence between the CBh promoter and bGHpolyA signal of a transient expression plasmid vector (SEQ ID NO: 95). In addition, a vector expressing the T8-2 guide RNA was created by subcloning a DNA sequence (sequence number 171) targeting the Tet operator sequence (TetO) or a DNA sequence (sequence number 172) targeting a sequence not present in the reporter vector or human genome onto the 3' side of the guide RNA scaffold sequence of the T8-2 guide RNA expression vector created in Example 10.

[0070] Example 18: Construction of eGFP expression reporter vector to evaluate transcription activation by dead T8-2-P300 core An eGFP expression reporter vector for evaluating transcriptional activation by dead T8-2-P300 core in human cultured cells was constructed as follows. The reporter vector is a pUC-based transient expression plasmid vector (synthesized by VectorBuilder) (SEQ ID NO: 173) that contains a TRE3G promoter with a repeat sequence consisting of seven Tet operator sequences (TetO), and a Kozak sequence, an eGFP CDS sequence, and an SV40 polyA signal on its 3' side.

[0071] Example 19: Evaluation of transcription activation ability of dead T8-2-P300 core using HEK293FT 300 ng of the dead T8-2-P300 core expression vector constructed in Example 17, 300 ng of a T8-2 guide RNA expression vector (sgRNA (TetO target)) having the TetO sequence of the eGFP reporter vector as its target sequence, or 300 ng of a reporter vector and a guide RNA expression vector having a sequence not present in the human genome as its target sequence (sgRNA (No target)), together with 300 ng of the eGFP expression reporter vector prepared in Example 18, were mixed with Lipofectamine 3000 reagent and transfected into 96-well cultured HEK293FT cells (up to 10 4 cells / well) and transfection was performed. Transfected HEK293FT cells were cultured in Dulbecco's modified Eagle's medium containing 10% fetal bovine serum at 37°C under 5% CO2 conditions. After 24 hours of culture, the amount of eGFP expression was measured using a fluorescence microscope to assess the transcriptional activation ability of the cells (Figure 13). The integrated value of eGFP fluorescence intensity in cultured cells transfected with the dead T8-2-P300 core expression vector and gRNA (TetO) was approximately 7.1-fold higher than that in cells transfected with gRNA (No target). These results indicate that the dead T8-2-P300 core fusion protein induces transcriptional activation of genes located near the target sequence by altering the epigenetic state.

Claims

1. (i) a TnpB-like RNA-guided nuclease consisting of the amino acid sequence of SEQ ID NO: 20 or 22, or a mutant protein thereof, wherein the mutant protein has at least 90% sequence identity to the amino acid sequence and has TnpB-like RNA-guided nuclease activity; (ii) a guide RNA comprising a guide sequence and a sequence including a gRNA scaffold sequence; Here, when a TnpB-like RNA-guided nuclease consisting of the amino acid sequence of SEQ ID NO: 20 or a mutant protein thereof is used, the gRNA scaffold sequence comprises the nucleic acid sequence of SEQ ID NO: 99 or a sequence having at least 90% identity thereto; A TnpB-like RNA-guided nuclease complex, wherein when a TnpB-like RNA-guided nuclease consisting of the amino acid sequence of SEQ ID NO: 22 or a mutant protein thereof is used, the gRNA scaffold sequence comprises the nucleic acid sequence of SEQ ID NO: 100 or a sequence having at least 90% identity to said sequence.

2. The TnpB-like RNA-guided nuclease complex of claim 1, wherein the TAM sequence of the TnpB-like RNA-guided nuclease is a sequence containing TCAT.

3. 3. The TnpB-like RNA-guided nuclease complex of claim 2, wherein the TAM sequence comprises TTCAT.

4. 2. The TnpB-like RNA-guided nuclease complex of claim 1, wherein the mutant protein has an amino acid substitution at at least one position selected from the group consisting of Q at position 284, K at position 89, and R at position 291.

5. In the mutant protein, If there is an amino acid substitution at position 284, it is Q284R or Q284K; If there is an amino acid substitution at position 89, it is K89R; If there is an amino acid substitution at position 291, it is R291K. The TnpB-like RNA-guided nuclease complex of claim 4.

6. The TnpB-like RNA-guided nuclease complex according to claim 1, wherein the TnpB-like RNA-guided nuclease is a TnpB-like RNA-guided nuclease consisting of the amino acid sequence of SEQ ID NO: 22 or a mutant protein thereof, and the gRNA scaffold sequence comprises a nucleic acid consisting of SEQ ID NO: 117 (gRNA1), SEQ ID NO: 118 (gRNA2), SEQ ID NO: 119 (gRNA3), or a base sequence having 90% sequence identity to said sequence.

7. A method for cleaving a target DNA using the TnpB-like RNA-guided nuclease complex according to any one of claims 1 to 6 (excluding cleavage methods in living human cells).

8. The cutting method comprises: a nucleic acid comprising an expression cassette encoding the TnpB-like RNA-guided nuclease; The cleavage method according to claim 7, comprising introducing into a cell a nucleic acid comprising an expression cassette encoding the guide RNA.

9. The cleavage method according to claim 7, which is used as a genome editing method including sequence-specific knockout and sequence-specific knock-in.

10. A genome editing kit comprising the TnpB-like RNA-guided nuclease complex according to any one of claims 1 to 6.

11. The genome editing kit of claim 10, wherein the guide RNA comprising the gRNA scaffold sequence comprises a guide sequence designed from the cleavage target sequence.

12. The genome editing kit of claim 10, wherein the TAM sequence of the TnpB-like RNA-guided nuclease is a sequence containing TCAT.

13. The genome editing kit according to claim 10, wherein the genome editing kit is used for sequence-specific knockout and sequence-specific knockin.

Citation Information

Patent Citations

  • Reprogrammable TNPB polypeptides and use thereof

    WO2022159892A1