Tnpb-like RNA guide nuclease complex
The novel TnpB-like RNA-guided nuclease complex addresses the limitations of current genome editing technologies by enabling efficient and specific DNA cleavage without the need for PAM sequences, enhancing targeting flexibility and precision.
Patent Information
- Application Number
- PCT/JP2024/045047
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-12-19
- Publication Date
- 2025-06-26
AI Technical Summary
Current genome editing technologies, such as CRISPR/Cas9 and CRISPR/Cas12a, require specific PAM sequences and have limitations in targeting efficiency and specificity.
A novel TnpB-like RNA-guided nuclease complex is developed, comprising a TnpB-like RNA-guided nuclease and a guide RNA with a specific sequence, which can cleave target DNA independently of PAM sequences, enhancing targeting flexibility and efficiency.
The TnpB-like RNA-guided nuclease complex achieves efficient and specific DNA cleavage, allowing for precise genome editing and transcription regulation, with improved targeting capabilities compared to existing technologies.
Smart Images

Figure JP2024045047_26062025_PF_FP_ABST
Abstract
Description
TnpB-like RNA-guided nuclease complex
[0001] The present invention relates to a novel TnpB-like RNA-guided nuclease complex, a technique for cleaving target DNA using the complex, a genome editing technique involving cleavage, and a technique for regulating transcription in target DNA using the TnpB-like RNA-guided nuclease complex.
[0002] Genome editing is a technique for introducing (editing) mutations at targeted locations in the genome of a target organism. Genome editing technology using CRISPR-Cas is based on the discovery of adaptive immune systems in bacteria and archaea against foreign viruses and plasmids, and is a technology that applies Cas, an RNA-guided endonuclease involved in such adaptive immune systems. DNA from foreign viruses and plasmids is incorporated into the CRISPR locus by the action of the Cas protein. Once incorporated into the CRISPR locus, a portion of the foreign DNA functions as a template for producing crRNA. The produced crRNA forms a complex with tracrRNA, functioning as a guide strand to recruit Cas protein to foreign DNA complementary to the crRNA, and cleaves the foreign DNA near the PAM sequence, thereby eliminating the foreign DNA and exerting immune function. Focusing on the ability to cleave a specific sequence, guide RNAs that exert the functions of crRNA and tracrRNA are designed to be complementary to the target DNA sequence, thereby inducing DNA cleavage at the desired site. Therefore, compared to existing mutagenesis methods such as radiation, the advantage of this method is the efficiency with which the desired strain can be obtained from an overwhelmingly smaller number of mutants, and it is used for a variety of purposes.
[0003] In typical genome editing technologies such as CRISPR / Cas9 and CRISPR / Cas12a, the sequence to be edited is primarily determined by a portion of the gRNA or crRNA that constitutes them, known as the spacer sequence. The gRNA or crRNA spacer sequence is typically used as a target sequence that can be arbitrarily modified, typically consisting of 20 to 24 bases. In addition to this target sequence, the nucleases Cas9 and Cas12a require a base sequence of approximately 2 to 4 bases, known as a PAM sequence, to function. Therefore, genome editing is performed by specifically recognizing a total base sequence of approximately 22 to 28 bases, consisting of the PAM sequence and the target sequence. The DNA sequence of the cleavage target must contain a PAM sequence, but the PAM sequence varies depending on the type of Cas protein family. By selecting a Cas protein that matches the target DNA sequence, genome editing of the desired sequence becomes possible.
[0004] It has been speculated that the RNA-guided endonuclease Cas protein originates from the IscB and TnpB proteins contained in transposons (Non-Patent Document 1: J. Bacteriol. 198, 797-807 (2015)). The IscB and TnpB proteins, which are nucleases contained in transposons, have been reported to function as RNA-guided endonucleases (Non-Patent Document 2: Science 374, 57-65 (2021); Non-Patent Document 3: Nature 599, 692-696 (2021)). It has also been reported that a complex containing a protein classified as TnpB, which contains a Ruv-C nuclease domain, and an ωRNA molecule can be used for genome editing (Patent Document 1: International Publication No. 2022 / 159892), and active research is being conducted on functional analysis of TnpB and identification of TnpB orthologs and their use as new genome editing tools (Non-Patent Document 4: The CRISPR Journal. Jun 2023 pp. 232-242, Non-Patent Document 5: Nat Biotechnol., 2023), Non-Patent Document 6: Nature 620, 660-668, 2023, Non-Patent Document 7: Science Advances Sep 2023 vol. 9 Issue 39, Non-Patent Document 8: Nucleic Acids Research, gkad1053,2023, Non-Patent Document 9: Proc Natl Acad Sci USA. Nov 2023). 28;120(48):e2308224120.).
[0005] International Publication No. 2022 / 159892
[0006] J. Bacteriol. (2015) 198, 797-807Science (2021)374, 57-65Nature (2021) 599, 692-696The CRISPR Journal.(2023) vol. 6, No. 3, p.232-242Nat Biotechnol., 2023,June 29, p.1-13Nature, 2023,620,660-668Science Advances Sep 2023 vol. 9 Issue 39, eadlk0171Nucleic Acids Research, gkad1053,2023Proc Natl Acad Sci USA. 2023 Nov 28;120(48):e2308224120.Int J Mol Sci., 2020 Apr. 25;21(9):3038. doi: 10.3390 / ijms21093038.
[0007] The aim is to identify and utilize novel RNA-guided nucleases that can be applied to genome editing technology.
[0008] The present inventors conducted extensive research to identify new proteins not annotated as TnpB that form RNA-guided endonuclease (hereinafter also referred to as RGN) complexes applicable to genome editing technology. As a result, they discovered a TnpB-like RNA-guided nuclease that is not classified as TnpB, leading to the present invention. The present invention therefore relates to the following: [1] (i) a TnpB-like RNA-guided nuclease having a length of 300 to 700 amino acids and comprising the following: YKGRTFNKMINNGSKGQYNxR-SxNxLKWRG (SEQ ID NO: 109) [wherein x represents any amino acid]; and (ii) a TnpB-like RNA-guided nuclease complex comprising a guide sequence and a gRNA scaffold sequence. [2] The TnpB-like RNA-guided nuclease complex according to Item 1, wherein the TnpB-like RNA-guided nuclease is a protein consisting of an amino acid sequence selected from the group consisting of SEQ ID NOs: 16, 20, 22, 24, 26, 40, 50, 52, 54, and 56. [3] The TnpB-like RNA-guided nuclease complex according to Item 1, wherein the TnpB-like RNA-guided nuclease is a mutant protein having a mutation involving an amino acid substitution with respect to the protein according to Item 2, and wherein the mutant protein has RNA-guided nuclease activity. [4] The TnpB-like RNA-guided nuclease complex according to Item 3, wherein the mutant protein consists of an amino acid sequence that has at least 75% sequence identity to the amino acid sequence of the TnpB-like RNA-guided nuclease complex according to Item 2. [5] The TnpB-like RNA-guided nuclease complex according to Item 3, wherein the mutant protein has structural homology to the original protein in the RuvC domain, and the structural homology is such that, when the original protein and the mutant protein are structurally aligned using PyMOL software, the C root mean square deviation (RMSD) between the backbone of the original protein and the backbone of the mutant protein is 1.5 or less and the TM value is 0.9 or more.[6] The TnpB-like RNA-guided nuclease complex according to any one of Items 1 to 5, wherein the TAM sequence of the TnpB-like RNA-guided nuclease is a sequence comprising TCAT. [7] The TnpB-like RNA-guided nuclease complex according to Item 6, wherein the TAM sequence comprises TTCAT. [8] The TnpB-like RNA-guided nuclease complex according to any one of Items 1 to 7, wherein the TnpB-like RNA-guided nuclease is a TnpB-like RNA-guided nuclease consisting of the amino acid sequence of SEQ ID NO: 22 or a mutant protein thereof, and the gRNA scaffold sequence comprises the sequence of SEQ ID NO: 58 or a sequence having at least 70% identity to said sequence. [9] The TnpB-like RNA-guided nuclease complex according to Item 8, wherein the mutant protein has an amino acid substitution at at least one position selected from the group consisting of Q at position 284, R at position 458, K at position 89, S at position 72 and / or S at position 75, R at position 291, S at position 201, L at position 278, and S at position 315.
[10] The TnpB-like RNA-guided nuclease complex according to Item 9, wherein the mutant protein has an amino acid substitution at position 284, which is Q284R or Q284K; an amino acid substitution at position 89, which is K89R; an amino acid substitution at position 72 or 75, which is S72H, S72N, S72Q, S72K, S72R, S75H, S75N, S75Q, S75K, or S75R; an amino acid substitution at position 291, which is R291E or R291K; an amino acid substitution at position 201, which is Q201H or Q201S; an amino acid substitution at position 278, which is L278K or L278R; and an amino acid substitution at position 315, which is S315A or S135V.
[11] The TnpB-like RNA-guided nuclease complex according to any one of Items 1 to 10, wherein the TnpB-like RNA-guided nuclease is a TnpB-like RNA-guided nuclease consisting of the amino acid sequence of SEQ ID NO: 22 or a mutant protein thereof, and the gRNA scaffold sequence comprises a nucleic acid consisting of SEQ ID NO: 117 (gRNA1), SEQ ID NO: 118 (gRNA2), SEQ ID NO: 119 (gRNA3), or a base sequence having 90% sequence identity to said sequences.
[12] A method for cleaving a target DNA using the TnpB-like RNA-guided nuclease complex according to any one of Items 1 to 11.
[13] The cleavage method according to Item 12, comprising introducing into a cell a nucleic acid comprising an expression cassette encoding the TnpB-like RNA-guided nuclease and a nucleic acid comprising an expression cassette encoding the guide RNA.
[14] The cleavage method according to Item 12, wherein the cleavage method is used as a genome editing method including sequence-specific knockout and sequence-specific knock-in.
[15] An inactivating mutation-containing TnpB-like RNA-guided nuclease complex comprising: (i) a TnpB-like RNA-guided nuclease of the TnpB-like RNA-guided nuclease complex according to any one of Items 1 to 11, the TnpB-like RNA-guided nuclease having an inactivating mutation and fused to a transcription regulatory domain; and (ii) a guide RNA comprising a sequence including a guide sequence and a gRNA scaffold sequence.
[16] A method for regulating transcription of target DNA using the TnpB-like RNA-guided nuclease complex containing an inactivating mutation according to Item 15.
[17] The method for regulating transcription of target DNA according to Item 15, comprising introducing into a cell an expression cassette encoding the TnpB-like RNA-guided nuclease containing an inactivating mutation and an expression cassette encoding the guide RNA.
[18] A genome editing kit comprising the TnpB-like RNA-guided nuclease complex according to any one of Items 1 to 11.
[19] The genome editing kit according to Item 18, wherein the guide RNA containing a gRNA scaffold sequence comprises a guide sequence designed from the cleavage target sequence.
[20] The genome editing kit according to Item 18, wherein the TAM sequence of the TnpB-like RNA-guided nuclease is a sequence containing TCAT.
[21] The genome editing kit according to any one of Items 18 to 20, wherein the genome editing kit is used for sequence-specific knockout and sequence-specific knockin.
[0009] According to the present invention, a novel RNA-guided endonuclease enables cleavage of target nucleic acids, and can be applied to genome editing technology.
[0010] Figure 1 shows the phylogenetic relationships of five T protein groups (clades 1 to 5). The phylogenetic tree was constructed using the neighbor-joining method using the protein sequences listed in Table 1. Figure 2 shows the 3'-terminal sequences of genes encoding TnpB-like RNA-guided nucleases belonging to each clade as a result of sequence analysis. Focusing on the 3'-terminal boundary of the region of high homology between each clade, this boundary was determined as the TEM sequence. Figure 3 shows the results of SDS-PAGE separation and Coomassie blue staining of the TnpB-like RNA-guided nuclease solution obtained by expressing and column-purifying recombinant His-MBP-TEV fusion TnpB-like RNA-guided nuclease in E. coli (A: T2, B: T3, C: T4, D: T5, E: T6, F: T7, G: T8-1, H: T8-2). Figure 4 shows the results of extracting RNA complexed with TnpB-like RNA-guided nuclease and analyzing it by next-generation sequencing. It was shown that the RNA complexed with TnpB-like RNA-guided nuclease partially overlaps with the open reading frame of the TnpB-like RNA-guided nuclease gene. Figure 5 shows the results of analyzing the cleaved plasmids obtained by applying a complex of TnpB-like RNA-guided nuclease and guide RNA to a plasmid library containing random TAM and target sequences. Plasmids containing TCAT (T3), TTAT (T5 and T7), and TTCAT (T6, T8-1, T8-2) were cleaved in the random sequences, and these sequences became TAM sequences. Figure 6 shows the results of examining the cleavage activity of double-stranded DNA containing different target sequences amplified from purified plasmids containing TAM sequences, in vitro, using a complex of guide RNA corresponding to the target sequence and TnpB-like RNA-guided nuclease. Figure 7 shows the results of examining the cleavage activity of purified plasmids containing different TAM sequences and different target sequences in vitro, where a complex of a guide RNA corresponding to the target sequence and a TnpB-like RNA-guided nuclease (T5 or T8-2) was allowed to act. The types of RNPs acted upon and the TAM sequences used are shown at the top. No RNP represents the negative control.Figure 8 shows the results of examining the cleavage activity of purified plasmids containing TAM sequences and target sequences for T5 and T8 in vitro, in which a complex of guide RNA corresponding to the target sequence and TnpB-like RNA-guided nuclease (T5 and T8-2) was allowed to act, and the temperature was varied. Figure 9 is a graph showing the genome editing efficiency when HEK293FT was transfected with an expression vector encoding T8-1 protein or T8-2 and an expression vector encoding guide RNA. The vertical axis shows the absolute value of genome editing efficiency, and the horizontal axis shows the conditions when T8-1 gRNA1-5 and T8-2 gRNA1-5 were transformed, respectively. Control shows the genome editing efficiency when no T8 expression or guide DNA expression plasmid was included. Figure 10 shows the genomic nucleic acid sequence (5' to the left, 3' to the right) near the cleavage site when HEK293FT is transfected with an expression vector encoding the T8-1 protein and an expression vector encoding a guide RNA. The sgRNA indicates the target sequence and recognizes the complementary strand of the indicated DNA. The TAM sequence is located 3' to the target sequence and is not included in this figure. The dashed line indicates the site of the predicted double-stranded cleavage. Figure 11 shows the editing efficiency when modified T8-2 is introduced with gRNA original, which has a full-length 152-nt T8-2 guide RNA scaffold sequence, or ddHDV, which removes only the HDV sequence; gRNA1, in which the T8-2 guide RNA scaffold sequence has been shortened to 131 nt; gRNA2, in which the guide RNA scaffold sequence has been shortened to 111 nt; gRNA3, in which the guide RNA scaffold sequence has been shortened to 100 nt; and gRNA4, in which the guide RNA has been shortened to 77 nt. Figure 12 shows the effect of point mutations in modified T8-2. Figure 13 shows the increase in target protein production by an inactivating mutation-containing TnpB-like RNA-guided nuclease complex, which contains an inactivating mutation-containing T8-2, a TnpB-like RNA-guided nuclease fused to a P300 core domain, and a guide RNA.
[0011] One aspect of the present invention relates to a TnpB-like RNA-guided nuclease complex, a kit for expressing the complex, a method for cleaving target DNA using the TnpB-like RNA-guided nuclease complex, and a genome editing method.
[0012] [TnpB-like RNA-guided nuclease complex] The TnpB-like RNA-guided nuclease complex of the present invention comprises: (1) a TnpB-like RNA-guided nuclease; and (2) a gRNA comprising a sequence including a guide sequence and a gRNA scaffold sequence. Such a complex may cleave a specific DNA in a test tube, or may be formed in a cell to cleave a specific site in DNA complementary to the guide RNA sequence. The complex can be formed by genetically introducing and expressing a vector, messenger RNA, or the guide RNA itself containing an expression cassette encoding (1) the TnpB-like RNA-guided nuclease and (2) the guide RNA into a cell. Alternatively, a complex formed in vitro may be introduced into a cell using techniques such as liposomes or electroporation.
[0013] TnpB-like RNA-guided nucleases TnpB-like RNA-guided nucleases (hereinafter also referred to as T proteins) that form TnpB-like RNA-guided nuclease complexes comprise one or more of the following features: (i) the following regular expression sequence: YKGRTF-[NS]-[KR]-[LM]-x-[AN]-xG-[AS]-[KR]-[GS]-QYx(2)-R-[AS]-x-[DN]-xLxWxG (SEQ ID NO: 1) (wherein, [NS] represents N or S, [KR] represents K or R, [LM] represents L or S, x represents any amino acid, [AN] represents A or N, [AS] represents A or S, [KR] represents K or R, [GS] represents G or S, (x(2) represents any two amino acids, [AS] represents A or S, and [DN] represents D or N); (ii) comprises at least one domain selected from the group consisting of RuvCI, bridge helix, RuvCII, RuvCIII, and Wedge; (iii) has a length of 300 to 700 residues, preferably 500 to 675 residues, and more preferably 550 to 650; and (iv) has RNA-guided nuclease activity.
[0014] The regular expression sequence in the present invention is a conserved sequence among TnpB-like RNA-guided nucleases of clades 1 to 4. The regular expression sequence is a 29-amino acid sequence contained in the RuvC domain (more specifically, RuvCIII) of TnpB-like RNA-guided nucleases. By performing a database search using this regular expression sequence, it is possible to specifically and completely identify T proteins.
[0015] More specifically, the TnpB-like RNA-guided nuclease relates to a TnpB-like RNA-guided nuclease having the following regular expression sequence: (i) YKGRTFNKMINNGSKGQYNxR-SxNxLKWRG (SEQ ID NO: 109) (wherein each x independently represents any amino acid) and having a length of 300 to 700 amino acids. Such a regular expression sequence refers to a TnpB-like RNA-guided nuclease included in Clade 1, excluding T11, T16, and T18. Specifically, it may refer to TnpB-like RNA-guided nucleases represented by T6, T8-1 to T8-4, T15, T17, and T21-1 to T21-4, as an example.
[0016] The bridge helix, RuvCI, RuvCII, and RuvCIII are each thought to be involved in the endonuclease domain, and Cas, IscB, and TnpB also have similar domains. On the other hand, the regular expression sequence in RuvCIII is shared only by TnpB-like RNA-guided nucleases and does not match that of Cas9, Cas12, IscB, and TnpB. The TnpB-like RNA-guided nuclease according to the present invention may contain all domains of the bridge helix, RuvCI, RuvCII, and RuvCIII in order to exert RNA-guided nuclease activity. Furthermore, it may be identified by including a consensus sequence in each domain.
[0017] The RNA-guided nuclease activity of TnpB-like RNA-guided nucleases recognizes a 3- to 7-base sequence, particularly a 4- or 5-base sequence, called a TAM (Transposon Associated Motif) sequence, and cleaves the DNA of the target sequence downstream. Therefore, the target sequence to which the guide RNA binds is designed to be adjacent to the TAM sequence, and the TnpB-like RNA-guided nuclease targets and cleaves the target nucleic acid containing the target sequence and the TAM sequence. While the TAM sequence may differ for each TnpB-like RNA-guided nuclease, the same TAM sequence can usually be used within the same clade (Figure 5). Furthermore, TAM sequences can be determined using methods well known in the art. For example, a random TAM library in which a random site of several bases, e.g., 7 bases, is linked to the target sequence is subjected to a complex of the guide RNA of the target sequence and the TnpB-like RNA-guided endonuclease, and cleavage is confirmed. As an example, typical TAM sequences recognized by TnpB-like RNA-guided nucleases of each clade are as follows:
[0018] TnpB-like RNA-guided nucleases can also be referred to as T proteins (e.g., designated T1-T21), and can be phylogenetically classified into clades 1 to 5. Phylogenetic classification can be performed using methods well known in the art. As an example, homologs of TnpB-like RNA-guided nucleases can be searched for using NCBI's PHI-BLAST with the regular expression of the amino acid sequence (SEQ ID NO: 1) or the regular expression of clade 1 (SEQ ID NO: 109) as a seed query, and the results can be used to search for and collect proteins with sequence homology using BLASTP software. An alignment can then be performed using ClustalW or similar software, and a phylogenetic tree can be created from the alignment scores. As another example, the amino acid sequence of the T8-2 protein (SEQ ID NO: 22) can be used as a seed query to detect non-redundant protein sequences using NCBI BLASTP, and the protein sequence can be determined based on phylogenetic analysis of the protein sequence using the neighbor-joining method. As yet another example, protein families can also be classified based on the RMSD score obtained using FoldSeek with a partial 3D structure of the RuvC III domain as a query. The present invention particularly relates to T proteins of clades 1 to 4. Proteins included in clade 1 include T6, T8-1 to T8-4, T11, T15, T16, T17, T18, and T21-1 to T21-4. In particular, this class relates to TnpB-like RNA-guided nucleases characterized by not including T11, T16, and T18, and designated T6, T8-1 to T8-4, T15, T17, and T21-1 to T21-4. Proteins included in clade 2 include T2, T3, and T19. Proteins included in clade 3 include T4-1, T4-2, and T5. Proteins included in clade 4 include T1, T13, and T14. Proteins included in clade 5 include T7, T9, T10, T12, and T20. Phylogenetically, clade 5 is most closely related to known TnpB, followed by clade 4, clade 3, clade 2, and clade 1, in that order.
[0019] The amino acid sequences and ORF nucleotide sequences of proteins belonging to clade 1 are as follows:
[0020] The amino acid sequences and ORF nucleotide sequences of proteins belonging to clade 2 are as follows:
[0021] The amino acid sequences and ORF nucleotide sequences of proteins belonging to clade 3 are as follows:
[0022] The amino acid sequences and ORF nucleotide sequences of proteins belonging to clade 4 are as follows:
[0023] The amino acid sequences and ORF nucleotide sequences of proteins belonging to clade 5 are as follows:
[0024] In one aspect of the present invention, the TnpB-like RNA-guided nuclease of the present invention relates to a protein having an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 49, 50, 52, 54, and 56, or a mutant protein thereof. A mutant protein refers to a protein having an amino acid sequence that has at least 50% sequence homology or identity to the amino acid sequence of the original protein and that has TnpB-like RNA-guided nuclease activity. Such muteins further comprise the following canonical sequence in the RuvC domain: YKGRTF-[NS]-[KR]-[LM]-x-[AN]-xG-[AS]-[KR]-[GS]-QYx(2)-R-[AS]-x-[DN]-xLxWxG (SEQ ID NO: 1) wherein, [NS] represents N or S, [KR] represents K or R, [LM" represents L or S, x represents any amino acid, [AN] represents A or N, [AS] represents A or S, [KR] represents K or R, [GS" represents G or S, x(2) represents two any amino acids, [AS] represents A or S, and [DN] represents D or N. and the size of the RuvC domain is specified as a full-length of 200 to 350, preferably 250 to 300. Furthermore, such mutant proteins have steric structural homology, more preferably structural homology in the RuvC domain.
[0025] The mutant protein may be specified by a homology and / or identity of 40% or more, 45% or more, 50% or more, 55% or more, 60% or more, 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 99% or more to the amino acid sequence of the original protein. In particular, it is particularly preferred that the mutant protein belongs to the same clade as the original protein. When belonging to the same clade, the homology or identity is usually preferably 50% or more, 60% or more, 70% or more, or 80% or more.
[0026] Examples of TnpB-like RNA-guided nucleases belonging to Clade 1 include the following: (1) a protein consisting of an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 16, 20, 22, 24, 26, 32, 40, 42, 46, 50, 52, 54, and 56, particularly a protein consisting of an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 16, 20, 22, 24, 26, 40, 50, 52, 54, and 56; or (2) A protein having TnpB-like RNA-guided nuclease activity, comprising an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 16, 20, 22, 24, 26, 32, 40, 42, 46, 50, 52, 54, and 56, particularly an amino acid sequence having at least 50%, 60%, 70%, 75%, 80%, or 90% sequence identity or homology to an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 16, 20, 22, 24, 26, 40, 50, 52, 54, and 56. The sequence identity can be appropriately selected so as not to include T11, T16, or T18.
[0027] TnpB-like RNA-guided nucleases belonging to Clade 2 include the following: (1) a protein consisting of an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 4, 6, 8, and 48, or (2) a protein comprising an amino acid sequence having at least 40%, 50%, or 60% sequence homology or identity to an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 4, 6, 8, and 48, wherein the sequence has TnpB-like RNA-guided nuclease activity.
[0028] TnpB-like RNA-guided nucleases belonging to Clade 3 include the following: (1) a protein consisting of an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 10, 12, and 14, or (2) a protein comprising an amino acid sequence having at least 80% sequence homology or identity to an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 10, 12, and 14, and having TnpB-like RNA-guided nuclease activity.
[0029] TnpB-like RNA-guided nucleases belonging to Clade 4 include the following: (1) a protein consisting of an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 2, 36, and 38, or (2) a protein comprising an amino acid sequence having at least 60% sequence homology or identity to an amino acid sequence selected from the group consisting of the amino acid sequences of SEQ ID NOs: 2, 36, and 38, and having TnpB-like RNA-guided nuclease activity.
[0030] In one embodiment, the TnpB-like RNA-guided nuclease of the present invention is a protein encoded by a nucleic acid sequence selected from the group consisting of the nucleic acid sequences of SEQ ID NOs: 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 51, 53, 55, and 57; or (2) A protein encoded by a nucleic acid sequence having at least 60% sequence identity to a nucleic acid sequence selected from the group consisting of the nucleic acid sequences of SEQ ID NOs: 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 51, 53, 55, and 57, wherein the protein has TnpB-like RNA-guided nuclease activity.
[0031] In this specification, the "homology" of two amino acid sequences refers to the ratio of identical or similar amino acid residues appearing at corresponding positions when the two amino acid sequences are aligned, and the "identity" of two amino acid sequences refers to the ratio of identical amino acid residues appearing at corresponding positions when the two amino acid sequences are aligned. The "homology" and "identity" of two amino acid sequences can be determined using, for example, the BLAST (Basic Local Alignment Search Tool) program (Altschul et al., J. Mol. Biol., (1990), 215(3):403-10).
[0032] The mutant protein is a protein that has structural homology in the RuvC domain of the original protein, and may also have structural homology to the entire length of the original protein. The location of the RuvC domain varies depending on the type of T protein. For example, in the case of the T8-2 protein, R302 to E580 corresponds to the RuvC domain. For other T proteins, the RuvC domain corresponding to such positions can be determined. Structural homology can be determined by the RMSD value and / or TM value when aligned with the three-dimensional structure of the original protein. The RMSD value refers to the root-mean-square deviation of atomic positions. For example, when the C root-mean-square deviation (RMSD) between the backbone of the original protein and the backbone of the mutant is 2.5 A or less, preferably 1.5 A or less, and more preferably 1 A or less, it can be said that there is a high probability that the mutant will exhibit an effect equivalent to that of the original protein. The RMSD value can be determined using software well known in the art, such as PyMOL software align. In addition to or instead of the RMSD value, the TM value (template modeling score) can also be used. Proteins with a TM value of 0.9 or greater, preferably 0.91 or greater, and more preferably 0.93 or greater can be said to have structural homology. Therefore, variants can be identified by RMSD values in addition to or instead of sequence identity or homology. Specifically, when structural alignment is performed using PyMOL software between a T8-2 protein belonging to clade 1 and a TnpB-like RNA-guided nuclease belonging to clades 1 to 4, the RMSD value is 1.9 or less and the TM value is 0.9 or greater. In yet another example, when structural alignment is performed using PyMOL software between a T8-2 protein belonging to clade 1 and a TnpB-like RNA-guided nuclease belonging to clade 1, the RMSD value is 1.5 or less and the TM value is 0.91 or greater.Since all proteins within the clade have TnpB-like RNA-guided nuclease activity, if the RMSD between the original protein and the mutant protein is 1.9 or less, preferably 1.5 or less, and more preferably 1 or less, and the TM value is 0.9 or more, preferably 0.91 or more, and more preferably 0.93 or more, the mutant protein is inferred to have equivalent activity to the original, and can be used in genome editing tools using the procedures described herein.
[0033] In another example, the three-dimensional structures and homology of the T8-2 protein belonging to Clade 1 are compared with those of known highly active TnpB RNA-guided nucleases. An example of a known highly active TnpB RNA-guided nuclease is DraTnpB-AI (SEQ ID NO: 174). Aligning the three-dimensional structures allows identification of amino acid positions that contribute to activity. In the amino acid sequence of the T8-2 protein (SEQ ID NO: 22), the amino acid positions corresponding to Q at position 284, R at position 458, K at position 89, S at position 72 and / or S at position 75, R at position 291, S at position 201, L at position 278, and S at position 315 may contribute to the activity of TnpB-like RNA-guided nucleases belonging to Clade 1. The corresponding amino acid positions can be determined by structural or amino acid sequence comparison. In one aspect, the present invention relates to an amino acid sequence of a TnpB-like RNA-guided nuclease belonging to Clade 1, comprising a substitution at at least one position or a combination thereof selected from the group consisting of Q at position 284, R at position 458, K at position 89, S at position 72 and / or S at position 75, R at position 291, S at position 201, L at position 278, and S at position 315 in the amino acid sequence of SEQ ID NO: 22. These mutations are particularly preferably a combination of two positions or a combination of three positions. For example, a combination of K at position 89 and Q at position 284, K at position 89 and R at position 291, Q at position 284 and R at position 291, or a combination of K at position 89, Q at position 284, and R at position 291 is preferred.
[0034] Such substitutions include Q284R or Q284K when there is a substitution at position 284 or a corresponding position, K89R when there is a substitution at position 89 or a corresponding position, S72H, S72N, S72Q, S72K, S72R, S75H, S75N, S75Q, S75K, or S75R when there is a substitution at position 72 or 75 or a corresponding position, R291E or R291K when there is a substitution at position 291 or a corresponding position, Q201H or Q201S when there is a substitution at position 201 or a corresponding position, L278K or L278R when there is a substitution at position 278 or a corresponding position, and S315A or S135V when there is a substitution at position 315 or a corresponding position. These substitutions can enhance TnpB-like RNA-guided nuclease activity. More specifically, from the viewpoint of achieving high genome editing efficiency, substitutions of K89R, Q284K, Q284R, R291K, or combinations thereof are preferred. For example, K89R and Q284K, K89R and R291K, Q284K and R291K, or K89R, Q284K and R291K are preferred (FIG. 12).
[0035] The TnpB-like RNA-guided nuclease or its mutant protein of the present invention has TnpB-like RNA-guided nuclease activity over a wide temperature range, for example, 20 to 50°C, but is particularly active at 30 to 45°C, and more preferably around 37°C (Figure 8). Typical TnpB proteins and TnpB-derived proteins often exhibit maximum activity at around 40 to 50°C, allowing for efficient use at lower temperatures. Because they are highly active not only around 37°C, the body temperature of warm-blooded animals, but also at 25°C, they are particularly suitable for genome editing in living organisms such as plants, algae, mollusks, and microorganisms.
[0036] [Guide RNA (gRNA)] In the present invention, the guide RNA is designed based on the DNA sequence of the cleavage target and the TnpB-like RNA-guided nuclease used, and is configured to include a guide sequence and a gRNA scaffold sequence. The cleavage target DNA is selected under the condition that it contains a TAM sequence, since it is cleaved near the TAM sequence. A 14-25 nt sequence adjacent to the 3' side of the TAM sequence can be selected as the guide sequence. The guide sequence may match the DNA sequence of the cleavage target, or it may contain one or several mismatch sequences. If the guide sequence does not perfectly match the target sequence, cleavage may occur due to off-target effects. The gRNA scaffold must have a sequence that can be used by the TnpB-like RNA-guided nuclease used. Typically, the 3'-end boundary of the locus (i.e., the boundary between the gRNA scaffold and the guide sequence) can be determined by comparing the sequences of each TnpB-like RNA-guided nuclease at the locus. The four bases at the 3' end of the gRNA scaffold sequence can be defined as a transposon endo motif (TEM). The gRNA scaffold sequence can be 70 to 200 nt. The guide sequence can be designed to be adjacent to the 3' end of the scaffold sequence, i.e., the TEM sequence, or any sequence of 20 to 100 nt can be inserted.
[0037] For TnpB-like RNA-guided nucleases within each clade, the gRNA scaffold sequences have high homology / identity. For example, within the same clade, gRNA scaffold sequences can be used if they have 70% or more, preferably 80% or more, and more preferably 90% or more identity / homology. Representative gRNA scaffold sequences for each grade are as follows: The underlined parts represent the TEM sequence. Therefore, the gRNA scaffold can be any of the sequences set forth in SEQ ID NOs: 58 to 62, or a sequence having 70% or more, preferably 80% or more, and more preferably 90% or more sequence identity to the sequence.
[0038] Regions of the gRNA scaffold sequence that are more important for activity can be identified based on sequence identity within the clade, and shortened sequences containing such important regions can also be created. As an example, shortened sequences of the T8-2 gRNA scaffold sequence include SEQ ID NO: 117 (gRNA1), SEQ ID NO: 118 (gRNA2), and SEQ ID NO: 119 (gRNA3). Alternatively, shortened mutant sequences that have at least 90%, more preferably at least 92%, at least 93%, 95%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 119 (gRNA3) and are capable of functioning as RNA scaffolds can also be selected.
[0039] A gRNA designed based on the TnpB-like RNA-guided nuclease used and the nucleic acid sequence of the cleavage target can be expressed using techniques well known in the art. As an example, a nucleic acid containing an expression cassette in which DNA containing a gRNA sequence is placed under the control of a promoter can be used. Such an expression cassette may contain a transcription termination sequence in the 3' region. The nucleic acid containing the expression cassette is incorporated into a vector or plasmid and introduced into a specific cell in vitro or to produce the gRNA.
[0040] [TnpB-like RNA-guided nucleases containing inactivating mutations] TnpB-like RNA-guided nucleases can be inactivated by substituting amino acid residues in their active domains. Examples of activation domains include D308, E489, and D575 in T8-2. Substituting the residues at these positions can inactivate TnpB-like RNA-guided nucleases. While inactivated TnpB-like RNA-guided nucleases are recruited to target sequences by the action of guide RNA, they do not exhibit RNA-guided nuclease activity. By fusing a functional domain to a TnpB-like RNA-guided nuclease containing an activating mutation, its function on genes in the target sequence can be regulated. The functional domain to be fused may be any domain; examples include transcriptional regulatory domains such as p300, DMNT3, KRAB, SRDX, VPR, and VP64. In another example, any functional domain, such as APOBEC, AID, TadA, Dda, FokI, M-MLV-RTase, or GFP, can be used, enabling not only epigenetic state but also base modification, reverse transcription, phosphate bond cleavage, and photomanipulation of fluorescent proteins. RNA-guided nucleases containing inactivating mutations fused to functional domains are also widely known, including CRISPR-Cas RNA-guided nucleases (Non-Patent Document 10: Int J Mol Sci., 2020 Apr 25;21(9):3038. doi: 10.3390 / ijms21093038), and similar functional regulation is possible with TnpB-like RNA-guided nucleases.
[0041] The TnpB-like RNA-guided nuclease containing an inactivating mutation fused to a functional domain forms a complex with a guide RNA consisting of a sequence including a guide sequence and a gRNA scaffold sequence, and exerts a function according to the functional domain. When a transcription-promoting domain, such as all or a portion of the functional domain of p300, VPR, or VP64, is fused as the functional domain, transcription in the region recruited by the guide RNA is promoted. On the other hand, when a transcription-repressing domain, such as all or a portion of the functional domain of DMNT3, KRAB, or SRDX, is fused as the functional domain, transcription in the region recruited by the guide RNA is suppressed. Therefore, the present invention may relate to a method for regulating transcription of target DNA using a TnpB-like RNA-guided nuclease containing an inactivating mutation fused to such a functional domain.
[0042] [Target-Specific Cleavage Method] Another aspect of the present invention relates to a method for cleaving target DNA using the TnpB-like RNA-guided nuclease complex of the present invention. More specifically, the method relates to a method for cleaving target DNA using a complex comprising: (i) the TnpB-like RNA-guided nuclease of the present invention; and (ii) a guide RNA having a sequence including a guide sequence and a gRNA scaffold sequence. The complex may be formed in vitro and introduced into cells using electroporation, liposomes, or the like to cleave the target DNA, or the target DNA may be cleaved in vitro. Alternatively, the complex may be incorporated into an expression cassette and introduced into cells using a vector, messenger RNA, or the guide RNA itself to express the TnpB-like RNA-guided nuclease and the guide RNA in the cells, thereby cleaving genomic DNA. Such a cleavage method can be performed using a kit for cleaving target nucleic acids that includes an expression cassette.
[0043] Gene introduction can be carried out using techniques known in the art. The introduction method is not particularly limited and can be selected appropriately depending on the type of substance to be introduced and the target of introduction. Introduction methods are broadly divided into direct methods and methods using viral vectors. Direct methods include electroporation, liposome methods, particle gun methods in which gold particles are injected together with the gene, and whisker methods. For methods using viral vectors, adenovirus, adeno-associated virus, lentivirus, Agrobacterium, tobacco mosaic virus (TMV), etc. can be used as vectors depending on the host species.
[0044] Target-specific cleavage occurring within cells can be repaired by genomic DNA repair mechanisms. Non-homologous end joining or homologous recombination occurs during this process, enabling gene knockout or knock-in. Therefore, the target-specific cleavage method of the present invention can also be referred to as a genome editing method. The genome editing method of the present invention may be performed on cultured cells, cultured tissues, or living cells. The target species is not particularly limited, but genome editing is possible for any bacterial, archaeal, or eukaryotic cell. Eukaryotic cells may include any cell, such as plant cells, insect cells, or animal cells. For example, genome editing is possible for any mammalian cell, and genome editing is possible for both human and non-human animal cells. The T protein of the present invention is characterized by a low optimum temperature. Therefore, genome editing can be performed at temperatures of 10 to 30°C.
[0045] [Target-Specific Protein Recruitment] Another aspect of the present invention relates to a method for recruiting a specific protein to a target DNA using a TnpB-like RNA-guided nuclease complex according to the present invention, the complex being engineered to not exhibit TnpB-like RNA-guided nuclease activity. A complex engineered to not exhibit TnpB-like RNA-guided nuclease activity may be obtained by inactivating the activity of the TnpB-like RNA-guided nuclease, or by adjusting the length of the guide RNA to prevent the TnpB-like protein from exhibiting DNA cleavage activity while still containing a guide RNA that contains a sequence long enough to specifically bind to the target. Such a guide RNA may be a guide sequence of a target sequence of about 10 bases. More specifically, the TnpB-like RNA-guided nuclease complex used in the method for recruiting a specific protein to target DNA comprises: (i) a TnpB-like RNA-guided nuclease according to the present invention in which the amino acid residues in the active center have been replaced with other amino acids; (ii) a guide RNA comprising a sequence including a guide sequence and a gRNA scaffold sequence; and (iii) a specific protein that binds to or interacts with the TnpB-like RNA-guided nuclease. By expressing or introducing such a complex into a cell, the complex is recruited near the site of the guide sequence, without cleavage, and the specific protein bound to the complex can be recruited. The active center of the mutant TnpB-like RNA-guided nuclease described in (i) can be a typical DED active center contained in the RuvC domain. In the case of T8-2, the activity of the TnpB-like guided nuclease can be inactivated by introducing mutations at D308, E489, and D575. The mutation can be appropriately selected within a range that can inactivate activity, and an example is a mutation to alanine. A binding site for other proteins, such as an MS2 sequence, can be inserted into the RNA scaffold sequence of the guide RNA described in (ii). The specific protein described in (iii) can be a protein that modifies DNA or epigenetic conditions, specifically, any protein commonly known as an epigenetic factor.More specifically, these include p300, DMNT3, KRAB, SRDX, VPR, and VP64. In addition, they not only affect epigenetic states but also base modifications, reverse transcription, cleavage of phosphate bonds, and binding of light-manipulation tools such as fluorescent proteins (specifically, APOBEC, AID, TadA, Dda, FokI, M-MLV-RTase, GFP, etc.), but are not limited to these. By recruiting such specific proteins to a target DNA region, the epigenetic state in the target DNA region can be changed.
[0046] [Kit] Another aspect of the present invention may relate to a kit for cleaving a target nucleic acid in a cell or a kit for genome editing. The kit comprises: (i) an expression cassette for the TnpB-like RNA-guided nuclease of the present invention; and (ii) an expression cassette for a guide RNA having a sequence including a gRNA scaffold sequence. The guide RNA expression cassette is provided so that a guide sequence can be inserted. A guide RNA expression cassette can be prepared by designing a cleavage target sequence and introducing a guide sequence. The expression cassette for the TnpB-like RNA-guided nuclease and the expression cassette for the guide RNA into which the guide sequence has been introduced are incorporated into, for example, a plasmid or vector, and then genetically introduced into cells.
[0047] An expression cassette typically contains a promoter sequence and can be prepared by placing a sequence encoding a TnpB-like RNA-guided nuclease or guide RNA under its control. The promoter can transiently or constitutively control the expression of downstream sequences. The expression cassettes may be contained on a single polynucleotide or on different polynucleotides. The expression cassette may also contain elements contributing to expression, such as a terminator, as well as elements necessary for the preparation of a plasmid or vector, such as a signal sequence, a tag sequence, a reporter sequence, a multicloning site, a drug resistance gene, and an origin of replication, as long as the activity of the TnpB-like RNA-guided nuclease is not impaired. To ensure that the expressed protein functions within the nucleus, it is desirable to add a nuclear localization signal to the signal sequence. One or more nuclear localization signals may be arranged in series. The nuclear localization signal can be selected appropriately depending on the organism. This allows the TnpB-like RNA-guided nuclease expressed in the cell to translocate into the nucleus, where it can cooperate with the guide RNA to cleave the target nucleic acid sequence.
[0048] The genome editing kit of the present invention may further include other materials, reagents, tools, etc. necessary for carrying out the genome editing method of the present invention, such as nucleic acid introduction reagents and buffer solutions, as needed. Other materials necessary for carrying out the genome editing method of the present invention include a donor polynucleotide in addition to an expression cassette for a TnpB-like RNA-guided nuclease and / or an expression cassette for a guide RNA. Introducing the donor polynucleotide into the nucleus enables knock-in of the donor polynucleotide at the cleavage site created by the CRISPR / Cas system. The donor polynucleotide causes homologous recombination at the cleavage site by locating sequences homologous to the 5' and 3' sequences of the introduced sequence, respectively.
[0049] All documents mentioned herein are incorporated by reference in their entirety.
[0050] The present invention will be described in more detail below with reference to examples. However, these examples are merely examples shown for the convenience of explanation, and the present invention is not limited to these examples in any sense.
[0051] Example 1: Identification of TnpB-like RNA-guided nuclease homologs. TnpB-like RNA-guided nuclease homologs were detected using NCBI BLASTP and TBLASTN against NCBI non-redundant protein sequences, using the amino acid sequence of the T8-2 protein (SEQ ID NO: 22) as a seed query. Similar searches were also performed on the MGNIFY database in addition to the NCBI database. The nucleotide sequences of the gene loci corresponding to each TnpB-like RNA-guided nuclease were obtained from the NCBI Genbank database. Based on phylogenetic tree analysis of protein sequences using the neighbor-joining method, TnpB-like RNA-guided nucleases were classified into five groups (clades 1 to 5) (Figure 1). The rightmost boundary of the locus (i.e., the boundary between the guide RNA (gRNA) scaffold and the guide sequence) was determined by aligning the genomic sequences of the loci encoding the T proteins listed on the left and comparing their 3' ends using ClustalW (https: / / www.genome.jp / tools-bin / clustalw). From the alignment, the guide-scaffold boundary was identified as the most downstream position where sequence conservation sharply decreases (Figure 2). A sequence motif consisting of four nucleotides located at the 3' end of the scaffold was defined as a transposon-encoded motif (TEM).
[0052] By aligning the amino acid sequences of TnpB-like RNA-guided nucleases in clades 1 to 4, the conserved sequence (canonical expression) among these sequences was determined: YKGRTF-[NS]-[KR]-[LM]-x-[AN]-xG-[AS]-[KR]-[GS]-QYx(2)-R-[AS]-x-[DN]-xLxWxG (SEQ ID NO: 1) (wherein, [NS] represents N or S, [KR] represents K or R, [LM] represents L or S, x represents any amino acid, [AN] represents A or N, [AS] represents A or S, [KR] represents K or R, [GS] represents G or S, x(2) represents two any amino acids, [AS] represents A or S, [DN] represents D or N.) By performing a database search using PHI-BLAST or the like for this regular expression sequence, it is possible to specifically identify all of the T proteins of clades 1 to 4.
[0053] Example 2: Construction of a TnpB-like RNA-guided nuclease expression vector. To produce recombinant TnpB-like RNA-guided nuclease in E. coli as an N-terminal 10x histidine-tagged maltose-binding protein-TEV protease cleavage site (His-MBP-TEV) fusion protein, a synthetic DNA fragment containing the His-MBP sequence was cloned between the NcoI and NdeI cleavage sites of pET28b using the NEBuilder HiFi DNA Assembly Kit (New England Biolabs). This was designated the pAN36 vector. The resulting plasmid was designated pAN36. The entire gene encoding the TnpB-like RNA-guided nuclease and the associated gRNA scaffold were synthesized by Integrated DNA Technology (IDT) (Table 8). For several synthetic TnpB-like RNA-guided nuclease genes, the 5' region of the coding sequence that did not overlap with the gRNA scaffold was codon-optimized for protein expression in E. coli. This DNA fragment was cloned into the pAN36 vector under a T7 promoter. In the resulting vector, the TnpB-like RNA-guided nuclease coding sequence was fused in-frame to the N-terminal His-MBP-TEV sequence. A description of the His-MBP-TEV-TnpB-like RNA-guided nuclease expression vector is shown in Table 8.
[0054] Example 3: Expression and purification of TnpB-like RNA-guided nuclease-RNA complex in E. coli To express the recombinant His-MBP-TEV fusion TnpB-like RNA-guided nuclease, E. coli Rosetta2(DE3)pLysS strain (Novagen) was transformed with the His-MBP-TEV-fusion TnpB-like RNA-guided nuclease expression vector. E. coli strains were cultured overnight at 37°C in Luria-Bertani (LB) medium supplemented with 50 μg / ml kanamycin and 34 μg / ml chloramphenicol. One liter of LB medium supplemented with 50 μg / ml kanamycin and 34 μg / ml chloramphenicol was inoculated with 10 ml of the overnight pre-culture, and the optical density (OD 600The cells were grown until the chromatin density (DDS) reached 0.6. Gene expression was then induced by adding 0.25 mM IPTG and grown at 18°C for 20 hours. The cells were harvested by centrifugation and stored at -70°C until use.
[0055] All subsequent purification steps were performed at 4°C. Cells were resuspended in buffer A (50 mM Tris-HCl, pH 8.0, 500 mM NaCl, 5% (v / v) glycerol, and 25 mM imidazole) supplemented with bovine DNAse I (Fujifilm Wako Pure Chemical Industries), chicken lysozyme (Fujifilm Wako Pure Chemical Industries), and protease inhibitors (phenylmethylsulfonyl fluoride and Roche cComplete ethylenediaminetetraacetic acid-free). After a 30-minute incubation, cells were disrupted by sonication (ULTRASONIC DISRUPTOR UD-211, TOMY) on ice. After centrifugation at 40,000 g for 30 min to remove cell debris, the supernatant was filtered through a 0.45 μm polyvinylidene fluoride (PVDF) membrane and batch-coupled to 1 ml of Ni-Sepharose 6 Fast Flow Resin (Cytiva) equilibrated for 1 h with buffer A. The resin was packed into an Econo-Pac chromatography column (Bio-Rad) and washed first with 14 ml of buffer B (50 mM Tris-HCl, pH 8.0, 1 M NaCl, 5% (v / v) glycerol, 25 mM imidazole) and then with 10 ml of buffer A. The bound protein was eluted with 3.5 ml of buffer C (50 mM Tris-HCl, pH 8.0, 1 M NaCl, 5% (v / v) glycerol, 300 mM imidazole). The proteins in the eluate were separated by SDS-PAGE on a 5-20% (w / v) polyacrylamide gel, and the gel was stained with Coomassie Brilliant Blue (Figure 3: SDS-PAGE gel). The peak fraction containing the fusion protein was transferred to a nuclease-free tube, flash-frozen in liquid nitrogen, and stored at -70°C until use. The resulting TnpB-like RNA-guided nuclease ribonucleoprotein (RNP) sample was used for nucleic acid extraction and dsDNA cleavage analysis. For in vitro double-stranded DNS cleavage analysis, TnpB-like RNA-guided nuclease RNP samples were further purified on a HiLoad 16 / 600 Superdex 200 pg column in buffer D [25 mM Tris-HCl pH 8.0, 150 mM NaCl, 5 mM MgCl, 1% (v / v) glycerol, 1 mM DTT].The pooled fractions were concentrated by ultrafiltration and stored at −80° C. until further use.
[0056] Example 4: Extraction and Analysis of gRNA Bound to TnpB-Like RNA-Guided Nucleases To extract RNA bound to TnpB-like RNA-guided nuclease RNPs, 180 μl of the peak fraction containing RNPs was vigorously mixed with 540 μl of TRI Reagent (Molecular Research Center, Inc.) and 108 μl of chloroform for 15 seconds, incubated at room temperature for 5 minutes, and then centrifuged at 12,000 g for 15 minutes at 4°C. The upper aqueous phase containing the RNA extracted from the RNPs was transferred to a new tube and mixed with 500 μl of 100% (v / v) ethanol. The sample was then loaded onto an RNA Clean & Concentrator-5 spin column (Zymo Research). After washing the spin column with wash buffer, 13.6 units of DNAse (QIAGEN) was applied to the spin column and incubated at room temperature for 15 minutes to remove residual DNA. After washing the spin column three times with buffer, the bound RNA was eluted with 15 μl of nuclease-free water. 500 ng of purified RNA was then used for RNA library preparation. RNA libraries were prepared using the SMARTer smRNA-Seq Kit for Illumina (TAKARA-BIO) according to the manufacturer's instructions. The resulting NGS adapter-ligated cDNA was amplified by eight cycles of PCR using full-length Illumina NGS indexing primers. Amplified DNA fragments of 200-500 bp were size-selected by 3% (w / v) agarose gel electrophoresis and gel-extracted and purified using the FastGene Gel / PCR Extraction Kit (FastGene). DNA fragments were pooled at equimolar ratios and adjusted to a 50 pM concentration using the Qubit dsDNA HS Assay Kit and Qubit 3.0 Fluorometer (Thermo Fisher Scientific Inc.). The pooled library was mixed with 0.25 volumes of 50 pM PhiX control v3 (Illumina Inc.) and then subjected to 2 × 150 bp paired-end sequencing on an iSeq100 (Illumina Inc.). Reads were adapter-trimmed and aligned to the template sequence using Bowtie2 software.The identified sequences of small RNAs bound to each TnpB-like RNA-guided nuclease were defined as the sequences of the corresponding gRNAs (Figure 4, a diagram of CDS and NGS reads merged). Information about the gRNAs is summarized in the table.
[0057] Example 5: Construction of 7N-TAM Library Plasmids The TAM sequences of TnpB-like RNA-guided nucleases were determined using a plasmid library containing seven randomized nucleotides (7N). To generate the 7N-TAM plasmid library, ssDNA containing seven randomized nucleotides (5'-AGCTATGACCATGATTACGAATTCNNNNNNNNNCTGCAGGAGCAAAGACC-3' (SEQ ID NO: 87)) was converted to dsDNA using the primer (5'-GTCTTTGCTCCTGCAG-3' (SEQ ID NO: 88)) with reverse transcriptase (ReverTra Ace, Takara). This dsDNA fragment was assembled with pUC18 containing an NGS adapter sequence using the NEBuilder HiFi DNA Assembly Cloning Kit (NEB) to generate the 7N-TAM plasmid library (Table 10, plasmid sequence). Escherichia coli DH5α cells were transformed with the reaction mixture and cultured at 37°C on LB plates supplemented with 100 μg / ml ampicillin. More than 100,000 colonies were washed from the plates, and the plasmid library was extracted using a NucleoBond Xtra Midiprep kit (Takara).
[0058] Example 6: Screening for TAM Sequences by In Vitro Transcription / Translation In vitro transcription / translation (IVTT) reactions were performed using the PUREfrex 2.0 Reconstituted Cell-Free Protein Synthesis Kit (GeneFrontier). According to the manufacturer's instructions, a DNA template encoding a T7 promoter-driven TnpB-like RNA-guided nuclease and a T7 promoter-driven gRNA scaffold with a spacer sequence targeting the TAM library (20 nt: 5'-GAATTCGTAATCATGGTCAT-3' (SEQ ID NO: 90)) was generated by PCR using the synthetic gene for TnpB-like RNA-guided nuclease as a template. The PCR primers used for DNA template preparation were Illumina's Nextera HT v2 dual index primers. IVTT reactions were performed in 20-25 μl reactions using 100 ng of TnpB-like RNA-guided nuclease template, 125 ng of the corresponding gRNA template with a spacer sequence, and 60 ng of TAM library plasmid. An IVTT reaction without gRNA DNA template served as a control. The reaction was incubated at 37°C for 4 hours, then quenched by adding 5 μl of 100 mg / ml RNase A (Nippon Gene) and incubated at 37°C for 30 minutes. After adding 5 μl of 1% SDS, the reaction was incubated at 50°C for 10 minutes to denature the protein. 4 units of protease K (FUJIFILM) were added and incubated at 37°C for 90 minutes to digest the protein. The undigested TAM library plasmid was recovered from the reaction using the FastGene Gel / PCR Extraction Kit and used as a template for subsequent PCR. Approximately 250-bp DNA fragments containing the TAM motif and target sequence were PCR-amplified for 30 cycles using KOD One PCR Master Mix (Toyobo) and Illumina full-length Nextera NGS indexing primers. The amplified DNA fragments were separated by 2.5% (w / v) agarose gel electrophoresis followed by gel extraction and purification. Preferentially depleted TAM motifs were identified by NGS amplicon sequencing of the TAM library plasmid as described above.TAM priorities were characterized from FASTQ files using a custom Python script. Briefly, 7-nucleotide TAMs were extracted and counted. TAM frequencies were normalized to the sequencing depth of each sample. Sequence displays were generated using WebLogo version 2.8.2 (http: / / weblogo.berkeley.edu / ) using TAMs with at least a 10-fold decline relative to the control (Figure 5).
[0059] Example 7: In vitro dsDNA cleavage analysis The 7N sequence was converted to the TAM sequence by site-directed mutagenesis, and a substrate plasmid containing the TAM sequence was synthesized. Substrate plasmids containing different target sequences (5'-GTTCTCCAGGCTGCTATCCTTAGCA-3' (SEQ ID NO: 91), 5'-GAATTCGTAATCATGGTCATAGCTG-3' (SEQ ID NO: 92)) were synthesized by site-directed mutagenesis. A 716-bp substrate dsDNA was amplified by PCR using the TAM sequence plasmid and primer 1 (5'-AAAGGGGATGTGCTGCAAGG-3' (SEQ ID NO: 93)) and primer 2 (5'-TATCTTTATAGTCCTGTCGG-3' (SEQ ID NO: 94)). The amplified dsDNA fragment (20 nM) was mixed with purified TnpB-like RNA-guided nuclease RNP (1, 2, or 4 μM) corresponding to the target sequence in 5 μl of buffer D and incubated for 2 hours. The reaction mixture was quenched by adding 2 μl of Proteinase K (Fujifilm Wako Pure Chemical Industries, Ltd.) and incubated at 60 °C for 5 minutes. The reaction products were analyzed using a MultiNA microchip electrophoresis system (SHIMADZU Inc.) (Figure 6).
[0060] Example 8: In vitro plasmid cleavage analysis. The 7N sequence was converted to a TAM sequence by site-directed mutagenesis, and substrate plasmids containing the TAM sequence were synthesized. Specifically, substrate plasmids were constructed with 5'-TTAT-3' or 5'-AACAT-3' for T5 and 5'-TTCAT-3' or 5'-AACAT-3' for T8-2, each of which had a sequence added to the 5' side of the target sequence 5'-GAATTCGTAATCATGGTCATAGCTG-3' (SEQ ID NO: 92). 50 ng of substrate plasmid was mixed with purified TnpB-like RNA-guided nuclease RNP (4 uM) corresponding to the target sequence in 5 μl of buffer D and incubated at 37°C for 2 hours. The reaction mixture was quenched by adding 2 μl of Proteinase K (Fujifilm Wako Pure Chemical Industries) and incubated at 60°C for 5 minutes. The reaction products were analyzed on a 1.0% (w / v) agarose gel (Figure 7), which showed a shift in mobility upon cleavage when the TAM sequence corresponded to the correct one.
[0061] Example 9: In Vitro Plasmid Cleavage Analysis (Confirmation of Temperature Dependence) 50 ng of the substrate plasmid used in Example 8 was mixed with purified TnpB-like RNA-guided nuclease RNP (4 μM) corresponding to the target sequence in 5 μl of Buffer D. The reaction temperature was varied from 20°C to 50°C in 5°C increments, and the mixture was incubated for 2 hours under each condition. The reaction mixture was quenched by adding 2 μl of Proteinase K (Fujifilm Wako Pure Chemical Industries, Ltd.) and then incubated at 60°C for 5 minutes. The reaction products were analyzed on a 1.0% (w / v) agarose gel (Figure 8). These results indicate that the TnpB RNA-guided nuclease used in this invention has a lower active temperature (around 37°C) compared to typical TnpB proteins and TnpB-derived genes, which reach their maximum activity around 40-50°C.
[0062] Example 10: Construction of T8 protein and gRNA expression vectors Plasmid vectors for expressing T8-1 or T8-2 protein and gRNA in human cultured cells were constructed as follows. The plasmid vector for expressing the T protein is a pUC-based transient expression plasmid vector (SEQ ID NO: 95) equipped with a CAG promoter, a multicloning site (MCS) in which SV40 NLS and Nuleoplasmin NLS are conferred to the C-terminus of the T protein, and a bGH polyA signal. The human codon-optimized T8-1 CDS sequence (SEQ ID NO: 96, synthesized by IDT) or the human codon-optimized T8-2 CDS sequence (SEQ ID NO: 97, synthesized by IDT) of the T protein was subcloned into the MCS. The plasmid vector for expressing gRNA is a pUC-based transient expression plasmid vector (SEQ ID NO: 98) that contains a human U6 promoter, a non-coding RNA cloning site, and a hammerhead ribozyme sequence on its 3' side. First, the gT8-1 gRNA scaffold sequence (SEQ ID NO: 99, synthesized by IDT) or the T8-2 gRNA scaffold sequence (SEQ ID NO: 100, synthesized by IDT) was introduced into the gRNA expression vector. Six sequences targeting the hAAVS1 gene (SEQ ID NOs: 101-106) were then subcloned.
[0063] Example 11: Evaluation of genome editing efficiency using HEK293FT Solutions containing 100 ng of the two types of T8 protein expression vectors (expressing T8-1 or T8-2) constructed in Example 10 and 50 ng of gRNA expression vectors (one of each of six types of gRNA expression vectors targeting the hAAVS1 gene) were mixed with Lipofectamine 3000 reagent and transfected into 96-well cultured HEK293FT cells (up to 10 4Cells / well) and transfection was performed with a total of 12 combinations. As a comparison, a condition in which a GFP protein expression plasmid (pEGFP-N1, Clonetech) was introduced was prepared (referred to as Control). After 3 days of culture, HEK293FT cell DNA was extracted with alkaline buffer (0.1 N NaOH) and roughly purified at 100 pg / μL as a template. KOD One PCR master Mix and the first round PCR primer set (gRNA1, gRNA2, gRNA3: SEQ ID NO: 110 and SEQ ID NO: 111, gRNA4, gRNA5, gRNA6: SEQ ID NO: 112 and SEQ ID NO: 113) were used to amplify the partial sequence of the hAAVS1 gene containing the genome editing target sequence under the following conditions. The target sequence was amplified. Similarly, a 20-fold dilution of the first-round PCR solution was used as a template, and a second-round PCR primer set, second-round PCR primers containing a Combinatorial Dual Index (CDI) sequence (5'-AATGATACGGCGACCACCGAGATCTACACNNNNNNNNNTCGTCGGCAGCGTC-3' (SEQ ID NO: 107) and 5'-CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTCTCGTGGGCTCGG-3' (SEQ ID NO: 108)) were used to carry out a PCR reaction under the following conditions, and a PCR amplification product having adapter sequences and index sequences attached to both ends for sequencing was obtained using an Illumina sequencing device. PCR conditions: 98°C / 2 minutes, 98°C / 10 seconds, 55°C / 5 seconds, 68°C / 3 seconds, 35 cycles (common to both the first and second rounds). The resulting PCR product was subjected to agarose gel electrophoresis, and the target nucleotide sequence was excised and purified. The concentration of the purified DNA was measured using a Qubit dsDNA HS Assay Kit and a Qubit 3.0 Fluorometer (Thermo Fisher Scientific, Waltham, MA, USA), and adjusted to 50 pM.The pooled library was mixed with 0.25 volumes of 50 pM PhiX control v3 (Illumina) and then subjected to 2 × 150-bp paired-end sequencing on an iSeq100 (Illumina, San Diego, CA, USA). The resulting Fastq files were analyzed using CRISPResso2 (Clement et al., 2019 Nature Biotechnology). If the T8 protein exhibits a cleavage pattern similar to that of TnpB, insertions and deletions within 10 bases downstream of the target sequence are likely genome editing induced by the T8 protein. Therefore, the mutation quantification window width was set to 10 bp from the predicted cleavage site (Figures 9 and 10). The "genome editing efficiency" was determined by measuring the percentage of all sequenced reads with insertions or deletions within the quantification window width. Figure 9 shows the results of evaluating genome editing efficiency using the above procedure for each of gRNA1-gRNA6, co-expressed with T8-1 or T8-2. The vertical axis represents genome editing efficiency (the percentage of all reads with insertion or deletion mutations in the quantification target region), and the horizontal axis represents the conditions for the combination of transformed vectors. While genome editing efficiency was nearly zero in the control, in both T8-1 and T8-2, co-expression with gRNA3 showed the highest genome editing efficiency (approximately 0.006). Next, gRNA5 was most efficient (approximately 0.0015). While T8-1 and T8-2 each showed similar genome editing efficiencies, gRNA1, gRNA2, and gRNA4 showed higher genome editing efficiency when co-introduced with T8-2 than when co-introduced with T8-1. Figure 10 shows a specific example of genome editing when a T8-2 protein expression vector and a gRNA3 expression vector were simultaneously introduced into HEK293FT cells and the genome sequence was analyzed. The target genome sequence is shown as Reference, the site where the gRNA is expected to bind is shown as sgRNA (rectangle), and the expected cleavage site is shown as a dashed line. Different sequences output by the sequencer were aligned, and the number of reads of the sequence (reads) and the percentage of total reads (%) are shown.Base substitutions are indicated in bold, and deletions are indicated by "-". When genome editing occurs, insertions or deletions are observed near the dashed line, which is the expected cleavage site, and multiple types of such reads (particularly deletions of 3 to 8 bases) have been detected at a rate of approximately 0.1%. From the above, it was found that in HEK293FT cells co-transfected with the T8-2 expression vector and the gRNA3 expression vector, genome editing, primarily deletions, occurs in the target genome region.
[0064] Example 12: Construction of modified T8-2 guide RNA expression vector Plasmid vectors for expressing T8-2 guide RNA in human cultured cells were constructed as follows. The plasmid vector for expressing gRNA is a pUC-based transient expression plasmid vector (SEQ ID NO: 98) that contains a human U6 promoter, a non-coding RNA cloning site, and a hammerhead ribozyme sequence on its 3' side. First, a 152-nt T8-2 guide RNA scaffold sequence (SEQ ID NO: 100, synthesized by IDT) was introduced into the gRNA expression vector. Furthermore, a sequence targeting the hAAVS1 gene (SEQ ID NO: 103) was subcloned. This guide RNA composed of the T8-2 guide RNA scaffold sequence, target sequence, and hammerhead ribozyme sequence was designated the gRNA original (SEQ ID NO: 114).
[0065] The modified T8-2 guide RNA expression vector was created as follows. A guide RNA expression vector (SEQ ID NO: 115) was created by removing only the hammerhead ribozyme sequence from the transient expression plasmid vector (SEQ ID NO: 98). A guide RNA sequence consisting only of the T8-2 guide RNA scaffold sequence (SEQ ID NO: 98) and the hAAVS1 gene target sequence (SEQ ID NO: 100) was cloned into the cloning site of this vector. This guide RNA sequence was designated ddHDV (SEQ ID NO: 116). Next, nucleotide sequences were gradually deleted from the 5' side of the guide RNA scaffold sequence within the ddHDV sequence, and four types of guide RNAs (synthesized by IDT) with shortened sequences were cloned: gRNA1 (guide RNA scaffold sequence length 131 nt, SEQ ID NO: 117), gRNA2 (guide RNA scaffold sequence length 111 nt, SEQ ID NO: 118), gRNA3 (RNA scaffold sequence length 100 nt, SEQ ID NO: 119), and gRNA4 (guide RNA scaffold sequence length 77 nt, SEQ ID NO: 120), to create a total of five modified T8-2 guide RNA expression vectors.
[0066] Example 13: Evaluation of genome editing efficiency by modified T8-2 guide RNA using HEK293FT Using the T8-2 protein expression vector constructed in Example 10 and the modified T8-2 guide RNA expression vector constructed in Example 12, the effect of modifying the guide RNA scaffold sequence on genome editing efficiency was evaluated using the method of Example 11 (Figure 11). For comparison, a condition in which a GFP protein expression plasmid (pEGFP-N1, Clonetech) was introduced was prepared (referred to as Control). In an experiment in which a T8-2 protein expression vector and a gRNA expression vector with a common target sequence but different gRNA coding regions were co-introduced, the genome editing efficiency of the hAAVS1 gene was compared. As a result, no insertion or deletion of bases into the hAAVS1 gene was observed in the Control. When gRNA original with a full-length 152nt T8-2 guide RNA scaffold sequence or ddHDV with only the HDV sequence removed was co-expressed with T8-2, the genome editing efficiency was approximately 0.7%. These results indicate that removing the hammerhead ribozyme sequence from the guide RNA does not affect genome editing efficiency. When gRNA1, in which the T8-2 guide RNA scaffold sequence was shortened to 131nt, was used, the genome editing efficiency was also approximately 0.7%, but when gRNA2, in which the guide RNA scaffold sequence was shortened to 111nt, was used, the genome editing efficiency was approximately 1.1%, and when gRNA3, in which the guide RNA scaffold sequence was shortened to 100nt, the genome editing efficiency improved to approximately 0.86%. However, no genome editing activity was observed with gRNA4, which was shortened to 77nt. These results indicate that deleting the 5' 21 nt of the original T8-2 guide RNA scaffold sequence improves genome editing activity, and furthermore, that the RNA sequences contained in the gRNA scaffolds of gRNA3 and gRNA4 play an important role in maintaining the DNA cleavage activity of T8-2.
[0067] Example 14: Construction of T8-2 mutant protein expression vectors Similar to Example 11, plasmid vectors were constructed to express 24 types of human codon-optimized T8-2 mutant proteins (SEQ ID NOs: 121-168) instead of the T8-2 protein.
[0068] Example 15: Evaluation of genome editing efficiency by T8-2 mutant protein using HEK293FT The genome editing efficiency of T8-2 mutant protein was evaluated in the same manner as in Example 12. The plasmid vector used for expressing gRNA was a plasmid vector containing SEQ ID NO: 118, which had the highest editing efficiency among the six gRNA vectors used in Example 13 (Figure 12). Figure 12 shows the results of evaluating genome editing efficiency using the above procedure under conditions in which unmutated T8-2 protein or each of 24 types of T8-2 mutant proteins was simultaneously co-expressed with a gRNA3 expression vector. The vertical axis shows genome editing efficiency (the proportion of reads with insertion or deletion mutations in the quantification target region out of all reads), and the horizontal axis shows the conditions of the transformed vector. The genome editing efficiency was nearly 0 in the control and approximately 0.3% when unmutated T8-2 protein was expressed. However, when the T8-2 mutant protein had the K89R mutation, the Q284K mutation, or the R291K mutation, the genome editing efficiency was approximately 1%. Furthermore, in T8-2 mutant proteins simultaneously carrying the K89R and Q284K mutations, or simultaneously carrying the K89R and R291K mutations, or simultaneously carrying the R291K and Q284K mutations, genome editing efficiencies of about 2% were observed in all cases. Furthermore, in T8-2 mutant proteins simultaneously carrying the K89R, R291K, and Q284K mutations, genome editing efficiencies of about 2.5% were observed. Therefore, it was concluded that the K89R, R291K, and Q284K mutations are mutations that increase the genome editing efficiency of the T8-2 protein, and that the combination of these mutations can produce T8-2 mutant proteins with improved genome editing.
[0069] Example 16: Construction of a fusion protein expression plasmid and gRNA expression vector composed of inactive T8-2 and the catalytic domain of the epigenetic factor histone acetyltransferase. In the known TnpB, a typical DED active center sequence contained in the endonuclease domain RuvC is known as the active center. It is known that substituting the N-terminal aspartic acid with alanine results in a loss of DNA cleavage activity. It is also known that expressing a fusion protein composed of a known inactive Cas9 and the catalytic domain of histone acetyltransferase (P300 core) and a guide RNA targeting a promoter sequence promotes acetylation of histone proteins present near the target sequence, thereby inducing transcriptional activation of nearby genes by changing the epigenetic state of the promoter region. With the aim of utilizing T8-2 in gene transcription activation technology, we aimed to create an artificial transcriptional activator composed of inactive T8-2 and P300 core, created by amino acid substitution in the DED active center sequence. A plasmid vector for expressing a fusion protein and gRNA consisting of inactive T8-2 and the catalytic domain P300core of histone acetyltransferase in human cultured cells was constructed as follows. A fusion protein (dead T8-2-P300core, SEQ ID NO: 169) consisting of an inactive T8-2 sequence in which the 308th aspartic acid residue constituting the DED active center sequence of T8-2 was replaced with alanine, a hemagglutinin tag on the C-terminal side of inactive T8-2, two SV40 NLS, and the P300core sequence of human histone acetyltransferase was encoded by a human codon-optimized CDS sequence (SEQ ID NO: 170, synthesized by IDT). The fusion protein expression vector was created by cloning the CDS sequence between the CBh promoter and bGHpolyA signal of a transient expression plasmid vector (SEQ ID NO: 95).Furthermore, a vector expressing the T8-2 guide RNA was prepared by subcloning a DNA sequence (SEQ ID NO: 171) targeting the Tet operator sequence (TetO) or a DNA sequence (SEQ ID NO: 172) targeting a sequence not present in the reporter vector or the human genome onto the 3' side of the guide RNA scaffold sequence of the T8-2 guide RNA expression vector prepared in Example 10.
[0070] Example 18: Construction of eGFP expression reporter vector for evaluating transcription activation by dead T8-2-P300 core An eGFP expression reporter vector for evaluating transcription activation by dead T8-2-P300 core in human cultured cells was constructed as follows. The reporter vector is a pUC-based transient expression plasmid vector (SEQ ID NO: 173, synthesized by VectorBuilder) that has a TRE3G promoter with a repeat sequence composed of seven Tet operator sequences (TetO), and on its 3' side has a Kozak sequence, an eGFP CDS sequence, and an SV40 polyA signal.
[0071] Example 19: Evaluation of transcription activation ability by dead T8-2-P300 core using HEK293FT 300 ng of the dead T8-2-P300 core expression vector constructed in Example 17, 300 ng of a T8-2 guide RNA expression vector (sgRNA (TetO target)) having the TetO sequence of the eGFP reporter vector as a target sequence or a reporter vector and a guide RNA expression vector (sgRNA (No target)) having a sequence not present in the human genome as a target sequence, and 300 ng of the eGFP expression reporter vector prepared in Example 18 were mixed with Lipofectamine 3000 reagent and transfected into 96-well cultured HEK293FT cells (up to 10 4cells / well) and transfection was performed. Transfected HEK293FT cells were cultured in Dulbecco's modified Eagle's medium containing 10% fetal bovine serum at 37 ° C and 5% CO2 conditions. After 24 hours of culture, the expression level of eGFP was measured using a fluorescence microscope to measure the integrated value of the fluorescence intensity generated by eGFP, and transcriptional activation ability was evaluated (Figure 13). The integrated value of eGFP fluorescence intensity in cultured cells transfected with the dead T8-2-P300 core expression vector and gRNA (TetO) was approximately 7.1 times higher than that of cells transfected with gRNA (No target). This result indicates that the dead T8-2-P300 core fusion protein induces transcriptional activation of genes located near the target sequence by changing the epigenetic state.
Claims
1. A TnpB-like RNA-guided nuclease complex comprising: (i) a TnpB-like RNA-guided nuclease comprising: YKGRTFNKMINNGSKGQYNxR-SxNxLKWRG (SEQ ID NO:109), where x represents any amino acid, and which is 300-700 amino acids in length; and (ii) a guide RNA comprising a sequence comprising a guide sequence and a gRNA scaffold sequence.
2. The TnpB-like RNA-guided nuclease complex of claim 1, wherein the TnpB-like RNA-guided nuclease is a protein consisting of an amino acid sequence selected from the group consisting of SEQ ID NOs: 16, 20, 22, 24, 26, 40, 50, 52, 54, and 56.
3. The TnpB-like RNA-guided nuclease complex described in claim 1, wherein the TnpB-like RNA-guided nuclease is a mutant protein having a mutation involving an amino acid substitution with respect to the protein described in claim 2, and has RNA-guided nuclease activity.
4. The TnpB-like RNA-guided nuclease complex of claim 3, wherein the mutant protein consists of an amino acid sequence having at least 75% sequence identity to the amino acid sequence of the TnpB-like RNA-guided nuclease complex of claim 2.
5. The TnpB-like RNA-guided nuclease complex of claim 3, wherein the mutant protein has structural homology to the original protein in the RuvC domain, and the structural homology is such that, when the original protein and the mutant protein are structurally aligned using PyMOL software, the C root mean square deviation (RMSD) between the backbone of the original protein and the backbone of the mutant is 1.5 or less and the TM value is 0.9 or more.
6. The TnpB-like RNA-guided nuclease complex of claim 1, wherein the TAM sequence of the TnpB-like RNA-guided nuclease is a sequence containing TCAT.
7. The TnpB-like RNA-guided nuclease complex of claim 6, wherein the TAM sequence comprises TTCAT.
8. The TnpB-like RNA-guided nuclease complex of claim 1, wherein the TnpB-like RNA-guided nuclease is a TnpB-like RNA-guided nuclease consisting of the amino acid sequence of SEQ ID NO: 22, or a mutant protein thereof, and the gRNA scaffold sequence comprises a sequence of SEQ ID NO: 58, or a sequence having at least 70% identity thereto.
9. The TnpB-like RNA-guided nuclease complex of claim 8, wherein the mutant protein has an amino acid substitution at at least one position selected from the group consisting of Q at position 284, R at position 458, K at position 89, S at position 72 and / or S at position 75, R at position 291, S at position 201, L at position 278, and S at position 315.
10. The TnpB-like RNA-guided nuclease complex of claim 9, wherein in the mutant protein, when there is an amino acid substitution at position 284, it is Q284R or Q284K; when there is an amino acid substitution at position 89, it is K89R; when there is an amino acid substitution at position 72 or 75, it is S72H, S72N, S72Q, S72K, S72R, S75H, S75N, S75Q, S75K, or S75R; when there is an amino acid substitution at position 291, it is R291E or R291K; when there is an amino acid substitution at position 201, it is Q201H or Q201S; when there is an amino acid substitution at position 278, it is L278K or L278R; and when there is an amino acid substitution at position 315, it is S315A or S135V.
11. The TnpB-like RNA-guided nuclease complex of claim 1, wherein the TnpB-like RNA-guided nuclease is a TnpB-like RNA-guided nuclease consisting of the amino acid sequence of SEQ ID NO: 22, or a mutant protein thereof, and the gRNA scaffold sequence comprises a nucleic acid consisting of SEQ ID NO: 117 (gRNA1), SEQ ID NO: 118 (gRNA2), SEQ ID NO: 119 (gRNA3), or a base sequence having 90% sequence identity to said sequence.
12. A method for cleaving a target DNA using the TnpB-like RNA-guided nuclease complex described in any one of claims 1 to 11.
13. The cleavage method according to claim 12, comprising introducing into a cell a nucleic acid comprising an expression cassette encoding the TnpB-like RNA-guided nuclease and a nucleic acid comprising an expression cassette encoding the guide RNA.
14. The cleavage method described in claim 12, which is used as a genome editing method including sequence-specific knockout and sequence-specific knock-in.
15. An inactivating mutation-containing TnpB-like RNA-guided nuclease complex comprising: (i) a TnpB-like RNA-guided nuclease having an inactivating mutation in the TnpB-like RNA-guided nuclease of the TnpB-like RNA-guided nuclease complex according to any one of claims 1 to 11, and fused to a transcriptional regulatory domain; and (ii) a guide RNA having a sequence including a guide sequence and a gRNA scaffold sequence.
16. A method for regulating transcription of a target DNA using a TnpB-like RNA-guided nuclease complex containing an inactivating mutation as described in claim 15.
17. The method for regulating transcription of a target DNA according to claim 15, comprising introducing into a cell an expression cassette encoding the TnpB-like RNA-guided nuclease containing an inactivating mutation and an expression cassette encoding the guide RNA.
18. A genome editing kit comprising the TnpB-like RNA-guided nuclease complex according to any one of claims 1 to 11.
19. The genome editing kit of claim 18, wherein the guide RNA comprising the gRNA scaffold sequence comprises a guide sequence designed from the cleavage target sequence.
20. The genome editing kit described in claim 18, wherein the TAM sequence of the TnpB-like RNA-guided nuclease is a sequence containing TCAT.
21. A genome editing kit described in any one of claims 18 to 20, wherein the genome editing kit is used for sequence-specific knockout and sequence-specific knock-in.
Citation Information
Patent Citations
Reprogrammable TNPB polypeptides and use thereof
WO2022159892A1