An isolated nuclease and its use
By developing an RNA-mediated endonuclease with a protein molecular weight less than spCas9 and high editing efficiency, the problems of difficulty in gene delivery and low editing efficiency in the prior art have been solved, and a more efficient gene editing effect has been achieved.
Patent Information
- Application Number
- CN202480001968.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-29
- Filing Date
- 2024-03-22
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-03-22
AI Technical Summary
The existing CRISPR/Cas9 gene editing technology is difficult to deliver the gene due to the excessive molecular weight of spCas9 protein, and the editing efficiency is not better than that of other systems, such as Cas12.
A new RNA-mediated endonuclease was developed with a protein molecular weight of less than spCas9, about one-third of it and has high gene editing efficiency. The amino acid sequence of the nuclease is specific, suitable for gene editing, and can bind to a specific guide RNA.
By delivering the new nuclease and guide RNA, the target gene that double-strand breaks to the host cell can be effectively introduced, solving the delivery difficulties caused by excessive molecular weight of spCas9 protein, and improving the efficiency of gene editing.
Smart Images

Figure CN119053698B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the priority of Chinese Patent Application No. 202310304837.4 filed on March 27, 2023 and PCT Application No. PCT / CN2023 / 135175 filed on November 29, 2023. The entire contents of these applications are hereby incorporated by reference in their entirety for all purposes. Technical field
[0003] This application relates to the field of molecular biology, and specifically relates to an isolated nuclease and its uses. This application also specifically relates to: a nucleic acid and nucleic acid construct encoding the nuclease, a guide RNA and its nucleic acid construct, and a composition, recombinant vector, recombinant host cell and kit containing the nuclease. This application also specifically relates to: a method for introducing a double - strand break into a target gene of a host cell, a method for deleting, replacing or inserting a target gene of a host cell, and a method for obtaining a host cell with a target gene deleted, replaced or inserted. This application also specifically relates to the uses of the nuclease, the nucleic acid and nucleic acid construct encoding the nuclease, the guide RNA and its nucleic acid construct, the composition, the recombinant vector, or the recombinant host cell in introducing a double - strand break into a target gene of a host cell, in deleting, replacing or inserting a target gene of a host cell, and in the preparation of drugs or preparations for gene therapy, cell therapy, genomic research, stem cell induction and post - induction differentiation. Background art
[0004] With the rapid development of modern biotechnology and the advent of the post - genomic era, people are entering the stage of rewriting and even redesigning genetic information from the stage of reading biological genetic DNA information. The discovery of the CRISPR / Cas9 technology has brought about a revolutionary breakthrough in gene editing technology. CRISPR / Cas9 is an RNA - mediated targeted gene - editing tool that can specifically recognize and cleave different endogenous DNA sequences by reprogramming sgRNA. Cas9 has two nuclease domains, RuvC and HNH, which are responsible for cleaving either strand of DNA. Mutation of any of these sites can convert Cas9 into a single - strand Cas9 nickase. Important new technologies related to Cas9, such as base editing and prime editing, are designed based on Cas9 nickase.
[0005] However, some drawbacks of CRISPR / Cas9 limit its applications: First, the CDS sequence of spCas9 has a length exceeding 4.1 Kb, which exceeds the maximum effective packaging capacity of adenovirus (AAV). Therefore, it is difficult for adenovirus-mediated gene delivery. Although the packaging capacity of lentivirus is stronger than that of AAV (with a carrying limit of about 9 kb), the protein of spCas9 still accounts for too large a proportion, resulting in limited room for subsequent modification. These drawbacks severely restrict the application of spCas9 in clinical medicine. Subsequently, smaller-molecular-weight CRISPR / Cas12 or 12f systems emerged, but the editing efficiency of proteins such as Cas12 is not better than that of spCas9. Therefore, spCas9 is still widely accepted and used at present. Second, the PAM sequence of spCas9 (which is the NGG sequence) is relatively simple and has a higher occurrence rate in the genome. Its advantage lies in the flexibility in reprogramming sgRNA to complete the recognition and cleavage of different DNA sequences. However, this flexibility also leads to off-target effects in less-than-ideal genome editing results.
[0006] Therefore, subsequently, gene editing technologies using RNA-mediated endonucleases emerged, namely the insertion sequences IscB and TnpB from the IS200 / IS605 family. They are widely distributed in microorganisms and have a more compact protein structure, with a size of about 400 aa, less than 1 / 3 of spCas9. Therefore, there is more room for modification in the application of the enzyme. TnpB mediates the cleavage of DNA next to the 5’TTGAT transposon-associated motif (TAM) through reRNA (right element RNA, derived from the RE element in the ISDra2 transposon), thereby causing break mutations in the DNA sequence on the genome. The DNA cleavage function of TnpB requires two conditions to be met simultaneously: (1) the TAM sequence; (2) the sequence located at the 3’ end of reRNA that matches the target gene. Since different nucleases can recognize different TAMs, the discovery of more highly active nuclease tools and the verification and detection of their functions can provide more, better, and more flexible options for the development of gene editing strategies.
[0007] It should be noted that the methods described in this section are not necessarily methods that have been previously envisioned or adopted. Unless otherwise specified, no method described in this section should be considered prior art solely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention
[0008] To solve the above problems, the present application aims to find an RNA-mediated endonuclease with an appropriate protein molecular weight and good gene editing effect, and provide more diverse and specific tools for gene editing.
[0009] The present application provides an isolated nuclease, wherein the nuclease comprises an amino acid sequence represented by the following formula:
[0010] (X1)(X2) a (X3)(X4)(X5) b (X6)(X7) c (X8)(X9) d (X 10 )(X 11 ) e (X 12 )(X 13 ) f (X 14 )(X 15 ) g (X 16 )
[0011] Wherein, a, b, c, d, e, f, g are the number of amino acids; (X1), (X3), (X4), (X6), (X8), (X 10 )、(X 12 )、(X 14 )、(X 16 ) are independently polar amino acids or aliphatic amino acids; (X2) is any amino acid, and a is 15 or 16; (X5) is any amino acid, and b is 2; (X7) is any amino acid, and c is 2, 3 or 4; (X9) is any amino acid, and d is 14, 15, 16, 17 or 18; (X 11 ) is any amino acid, and e is 1 or 2; (X 13 ) is any amino acid, and f is 6; and (X 15 ) is any amino acid, and g is 5.
[0012] According to embodiments of the present application, an isolated nuclease can be provided, wherein the nuclease has a nuclease sequence selected from the following (i) or a variant sequence of the foregoing nuclease having nuclease activity in (ii)-(iv): (i) at least one of the amino acid sequences shown in any one of SEQ ID NOs: 1-197; (ii) at least one of the sequences obtained by deleting, substituting, inserting or mutating 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids in the amino acid sequence shown in any one of SEQ ID NOs: 1-197; (iii) at least one of the amino acid sequences having at least 70%, 80%, 90%, 95% or 99% identity with the amino acid sequence shown in any one of SEQ ID NOs: 1-197; and (iv) at least one of the sequences obtained by further fusing other sequences to the amino acid sequence shown in any one of SEQ ID NOs: 1-197.
[0013] According to embodiments of the present application, a guide RNA can be provided, wherein the guide RNA contains a reRNA, the reRNA contains a nucleotide sequence shown in any one of SEQ ID NOs: 198-394 or a variant thereof, and the guide RNA can bind to a specific nuclease.
[0014] According to embodiments of the present application, a nucleic acid can be provided, wherein the nucleic acid encodes the nuclease described in the present application and / or the guide RNA described in the present application.
[0015] According to embodiments of the present application, a nucleic acid construct can be provided, which contains the nucleic acid described in the present application and also contains a promoter.
[0016] According to embodiments of the present application, a composition can be provided, wherein the composition contains: an IS200 / IS605 family nuclease or a functional fragment thereof, or a nucleic acid containing a coding for an IS200 / IS605 family nuclease or a functional fragment thereof, the nuclease or its functional fragment having endonuclease activity; and a guide RNA, or a nucleic acid containing a coding for a guide RNA, the guide RNA being capable of binding to a specific nuclease.
[0017] According to embodiments of the present application, a recombinant vector can be provided, wherein the recombinant vector contains a nucleic acid encoding the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, or the composition described in the present application.
[0018] According to embodiments of the present application, a recombinant host cell can be provided, wherein the recombinant host cell contains the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, the composition described in the present application, or the recombinant vector described in the present application.
[0019] According to an embodiment of the present application, a method for introducing a double-strand break into a target gene of a host cell can be provided, wherein the method comprises: delivering the nuclease, the guide RNA, the nucleic acid, the nucleic acid construct, the composition, or the recombinant vector described in the present application into the host cell.
[0020] According to an embodiment of the present application, a method for deleting, replacing, or inserting a target gene of a host cell can be provided, wherein the method comprises: delivering the nuclease, the guide RNA, the nucleic acid, the nucleic acid construct, the composition, or the recombinant vector described in the present application into the host cell.
[0021] According to an embodiment of the present application, a method for obtaining a host cell with a target gene deleted, replaced, or inserted can be provided, wherein the method comprises: delivering the nuclease, the guide RNA, the nucleic acid, the nucleic acid construct, the composition, or the recombinant vector described in the present application into the host cell.
[0022] According to an embodiment of the present application, uses of the nuclease, the guide RNA, the nucleic acid, the nucleic acid construct, the composition, the recombinant vector, or the recombinant host cell described in the present application in introducing a double-strand break into a target gene of a host cell can be provided.
[0023] According to an embodiment of the present application, uses of the nuclease, the guide RNA, the nucleic acid, the nucleic acid construct, the composition, the recombinant vector, or the recombinant host cell described in the present application in deleting, replacing, or inserting a target gene of a host cell can be provided.
[0024] According to an embodiment of the present application, uses of the nuclease, the guide RNA, the nucleic acid, the nucleic acid construct, the composition, the recombinant vector, or the recombinant host cell described in the present application in preparing drugs or preparations for gene therapy, cell therapy, genomic research, stem cell induction, and post-induction differentiation can be provided.
[0025] According to an embodiment of the present application, a kit can be provided, wherein the kit contains the nuclease, the guide RNA, the nucleic acid, the nucleic acid construct, the composition, the recombinant vector, or the recombinant host cell described in the present application.
[0026] The protein molecular weight of the nuclease described in this application is much smaller than that of spCas9, less than about one-third of it, providing more possibilities for various in vivo deliveries in subsequent gene therapies, and can solve the delivery difficulties brought by the protein size problem of spCas9. At the same time, compared with asCas12, which also has a relatively small protein molecular weight, this nuclease has a higher gene editing efficiency, thus providing the possibility of becoming a new gene editing application tool. In addition, since different nucleases can recognize different transposon-related motifs, the novel nuclease discovered in this application brings more choices for subsequent application scenarios of different scales.
[0027] It should be understood that the content described in this part is not intended to identify the key or important features of the examples of this application, nor is it used to limit the scope of this application. Other features of this application will become easily understood through the following description. Brief Description of the Drawings
[0028] The drawings exemplarily show the embodiments and form a part of the description. Together with the written description of the description, they are used to explain the exemplary embodiments of the embodiments. The shown embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0029] Figure 1 Shows a schematic diagram of the RGS dual-fluorescence surrogate reporting system in Example 1.
[0030] Figure 2 Shows a flow cytometry plot, where the percentage of mRFP+eGFP+ cells is presented in the upper right Q2 gate.
[0031] Figure 3Shows TP_A_1, TP_A_2, TP_A_8, TP_A_12, TP_A_18, TP_B_18, TP_B_41, TP_B_46, TP_B_70, TP_B_71, TP_B_72, TP_B_73, TP_C_23, TP_C_67, TP_C_70, TP_C_74, TP_D_1, TP_D_3, TP_D_4, TP_D_8, TP_D_17, TP_D_18, TP_D_23, TP_D_24, TP_D_25, TP_D_27, TP_D_30, TP_D_32, TP_D_40, TP_D_43, TP_D_51, TP_D_59, TP_D_61, TP_D_66, TP_D_67, TP_D_71, TP_D_72, TP_D_73, TP_E_2, TP_E_15, TP_E_17, TP_E_48, TP_F_56, TP_F_71, TP_F_77, TP_F_80, TP_F_83, TP_F_85, TP_G_14, TP_G_19, TP_G_20, TP_G_24, TP_G_43, TP_G_52, TP_G_53, TP_G_61, TP_G_66, TP_G_72, TP_G_75, TP_G_83, TP_G_84, TP_H_1, TP_H_3, TP_H_4, TP_H_5, TP_H_6, TP_H_9, TP_H_11, TP_H_12, TP_H_13, TP_H_15, TP_H_18, TP_H_19, TP_H_20, TP_H_21, TP_H_23, TP_H_24, TP_H_30, TP_H_31, TP_H_32, TP_H_34, TP_H_38, TP_H_39, TP_H_40, TP_H_43, TP_I_1, TP_I_2, TP_I_3, TP_I_4, TP_I_5, TP_I_6, TP_I_7, TP_I_8, TP_I_9, TP_I_10, TP_I_11, TP_I_12, TP_I_13, TP_I_15, TP_I_16, TP_I_17, TP_I_18, TP_I_19, TP_I_20, TP_I_21, TP_I_22, TP_I_24, TP_I_25, TP_I_26, TP_I_29, TP_I_31, TP_I_35, TP_I_37, TP_I_38, TP_I_40, TP_I_41, TP_I_44, TP_I_45, TP_I_46, TP_I_47, TP_I_48, TP_I_49, TP_I_50, TP_I_51, TP_I_52, TP_I_53, TP_I_55TP_I_56, TP_I_58, TP_I_59, TP_I_61, TP_I_62, TP_I_64, TP_I_65, TP_I_66, TP_I_67 , TP_I_70, TP_I_71, TP_I_76, TP_I_77, TP_I_79, TP_I_80, TP_I_82, TP_I_84, TP_I_8 5. TP_I_86, TP_I_87, TP_L_1, TP_L_4, TP_L_5, TP_L_8, TP_L_9, TP_L_10, TP_L_11, TP _L_12, TP_L_15, TP_L_16, TP_L_17, TP_L_21, TP_L_22, TP_L_24, TP_L_25, TP_L_26, TP GFP expression levels of all active nuclease candidates: TP_L_27, TP_L_28, TP_L_31, TP_L_32, TP_L_34, TP_L_36, TP_L_37, TP_L_39, TP_M_1, TP_M_3, TP_M_7, TP_M_11, TP_M_14, TP_M_17, TP_M_19, TP_M_20, TP_M_24, TP_M_31, TP_M_32, TP_M_33, TP_M_34, TP_M_35, TP_M_37, TP_M_40, TP_M_41, TP_M_43, TP_M_46, TP_M_49, TP_M_58, TP_M_65, TP_M_66, TP_M_67, TP_M_70, and TP_M_78. All results were quantified by flow cytometry assay as previously described in Example 2.
[0032] Figure 4 , Figure 5 , Figure 6 , Figure 7 and Figure 8 yes Figure 3 A partial enlarged view of .
[0033] Figure 9 The endogenous editing efficiency (quantified by the proportion of reads with insertions or deletions at the target site) of TP_C_23, TP_D_51, TP_D_67, TP_E_15, TP_F_85, TP_G_24, TP_H_5, TP_H_6, TP_H_9, TP_H_11, TP_H_24, TP_H_30, TP_H_32, TP_H_34, TP_H_38, TP_I_1, TP_I_5, TP_I_6, TP_I_12, TP_I_15, TP_I_18, TP_I_20, TP_I_38, TP_I_49, TP_I_64 and TP_I_79 in Example 3 is shown.
[0034] Figure 10Shows TP_A_1, TP_A_2, TP_A_8, TP_A_12, TP_A_18, TP_B_18, TP_B_41, TP_B_46, TP_B_70, TP_B_71, TP_B_72, TP_B_73, TP_C_23, TP_C_67, TP_C_70, TP_C_74, TP_D_1, TP_D_3, TP_D_4, TP_D_8, TP_D_17, TP_D_18, TP_D_23, TP_D_24, TP_D_25, TP_D_27, TP_D_30, TP_D_32, TP_D_40, TP_D_43, TP_D_51, TP_D_59, TP_D_61, TP_D_66, TP_D_67, TP_D_71, TP_D_72, TP_D_73, TP_E_2, TP_E_15, TP_E_17, TP_E_48, TP_F_56, TP_F_71, TP_F_77, TP_F_80, TP_F_83, TP_F_85, TP_G_14, TP_G_19, TP_G_20, TP_G_24, TP_G_43, TP_G_52, TP_G_53, TP_G_61, TP_G_66, TP_G_72, TP_G_75, TP_G_83, TP_G_84, TP_H_1, TP_H_3, TP_H_4, TP_H_5, TP_H_6, TP_H_9, TP_H_11, TP_H_12, TP_H_13, TP_H_15, TP_H_18, TP_H_19, TP_H_20, TP_H_21, TP_H_23, TP_H_24, TP_H_30, TP_H_31, TP_H_32, TP_H_34, TP_H_38, TP_H_39, TP_H_40, TP_H_43, TP_I_1, TP_I_2, TP_I_3, TP_I_4, TP_I_5, TP_I_6, TP_I_7, TP_I_8, TP_I_9, TP_I_10, TP_I_11, TP_I_12, TP_I_13, TP_I_15, TP_I_16, TP_I_17, TP_I_18, TP_I_19, TP_I_20, TP_I_21, TP_I_22, TP_I_24, TP_I_25, TP_I_26, TP_I_29, TP_I_31, TP_I_35, TP_I_37, TP_I_38, TP_I_40, TP_I_41, TP_I_44, TP_I_45, TP_I_46, TP_I_47, TP_I_48, TP_I_49, TP_I_50, TP_I_51, TP_I_52, TP_I_53 based on protein sequences in Example 2Phylogenetic trees of TP_I_55, TP_I_56, TP_I_58, TP_I_59, TP_I_61, TP_I_62, TP_I_64, TP_I_65, TP_I_66, TP_I_67, TP_I_70, TP_I_71, TP_I_76, TP_I_77, TP_I_79, TP_I_80, TP_I_82, TP_I_84, TP_I_85, TP_I_86, TP_I_87, TP_L_1, TP_L_4, TP_L_5, TP_L_8, TP_L_9, TP_L_10, TP_L_11, TP_L_12, TP_L_15, TP_L_16, TP_L_17, TP_L_21, TP_L_22, TP_L_24, TP_L_25, TP_L_26, TP_L_27, TP_L_28, TP_L_31, TP_L_32, TP_L_34, TP_L_36, TP_L_37, TP_L_39, TP_M_1, TP_M_3, TP_M_7, TP_M_11, TP_M_14, TP_M_17, TP_M_19, TP_M_20, TP_M_24, TP_M_31, TP_M_32, TP_M_33, TP_M_34, TP_M_35, TP_M_37, TP_M_40, TP_M_41, TP_M_43, TP_M_46, TP_M_49, TP_M_58, TP_M_65, TP_M_66, TP_M_67, TP_M_70, TP_M_78 and ISDra2.
[0035] Figure 11 and Figure 12 is Figure 10 a partial enlarged view of.
[0036] Figure 13 The schematic diagram shown shows the element order of the reporter vector in Example 4.
[0037] Figure 14 The schematic diagram shown shows how the reporter vector in Example 4 works.
[0038] Figure 15 、 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, and 31 show TP_A_1, TP_A_2, TP_A_8, TP_A_12, TP_A_18, TP_B_18, TP_B_41, TP_B_46, TP_B_70, TP_B_71, TP_B_72, TP_B_73, TP_C_23, TP_C_67, TP_C_70, TP_C_74, TP_D_1, TP_D_3, TP_D_4, TP_D_8, TP_D_17, TP_D_18, TP_D_23, TP_D_24, TP_D_25, TP_D_27, TP_D_30, TP_D_32, TP_D_40, TP_D_43, TP_D_51, TP_D_59, TP_D_61, TP_D_66, TP_D_67, TP_D_71, TP_D_72, TP_D_73, TP_E_2, TP_E_15, TP_E_17, TP_E_48, TP_F_56, TP_F_71, TP_F_77, TP_F_80, TP_F_83, TP_F_85, TP_G_14, TP_G_19, TP_G_20, TP_G_24, TP_G_43, TP_G_52, TP_G_53, TP_G_61, TP_G_66, TP_G_72, TP_G_75, TP_G_83, TP_G_84, TP_H_1, TP_H_3, TP_H_4, TP_H_5, TP_H_6, TP_H_9, TP_H_11, TP_H_12, TP_H_13, TP_H_15, TP_H_18, TP_H_19, TP_H_20, TP_H_21, TP_H_23, TP_H_24, TP_H_30, TP_H_31, TP_H_32, TP_H_34, TP_H_38, TP_H_39, TP_H_40, TP_H_43, TP_I_1, TP_I_2, TP_I_3, TP_I_4, TP_I_5, TP_I_6, TP_I_7, TP_I_8, TP_I_9, TP_I_10, TP_I_11, TP_I_12, TP_I_13, TP_I_15, TP_I_16, TP_I_17, TP_I_18, TP_I_19, TP_I_20, TP_I_21, TP_I_22, TP_I_24, TP_I_25, TP_I_26, TP_I_29, TP_I_31, TP_I_35, TP_I_37, TP_I_38, TP_I_40, TP_I_41, TP_I_44, TP_I_45, TP_I_46, TP_I_47, TP_I_48 in Example 4.The YFP expression levels of all active nuclease candidates of TP_I_49, TP_I_50, TP_I_51, TP_I_52, TP_I_53, TP_I_55, TP_I_56, TP_I_58, TP_I_59, TP_I_61, TP_I_62, TP_I_64, TP_I_65, TP_I_66, TP_I_67, TP_I_70, TP_I_71, TP_I_76, TP_I_77, TP_I_79, TP_I_80, TP_I_82, TP_I_84, TP_I_85, TP_I_86, TP_I_87, TP_L_1, TP_L_4, TP_L_5, TP_L_8, TP_L_9, TP_L_10, TP_L_11, TP_L_12, TP_L_15, TP_L_16, TP_L_17, TP_L_21, TP_L_22, TP_L_24, TP_L_25, TP_L_26, TP_L_27, TP_L_28, TP_L_31, TP_L_32, TP_L_34, TP_L_36, TP_L_37, TP_L_39, TP_M_1, TP_M_3, TP_M_7, TP_M_11, TP_M_14, TP_M_17, TP_M_19, TP_M_20, TP_M_24, TP_M_31, TP_M_32, TP_M_33, TP_M_34, TP_M_35, TP_M_37, TP_M_40, TP_M_41, TP_M_43, TP_M_46, TP_M_49, TP_M_58, TP_M_65, TP_M_66, TP_M_67, TP_M_70 and TP_M_78. Detailed implementation manners
[0039] Unless otherwise specified or inconsistent with the context, the terms or expressions used herein shall be read in combination with the entire content of this disclosure and as understood by those of ordinary skill in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art.
[0040] In this application, the terms "nucleic acid" and "polynucleotide" are used interchangeably and refer to polymeric forms of nucleotides of any length, including deoxyribonucleotides, ribonucleotides, combinations thereof, and analogs thereof.
[0041] In this application, the terms "polypeptide" and "peptide" are used interchangeably and refer to polymers of amino acids of any length. Thus, polypeptides, oligopeptides, proteins, antibodies, and enzymes are all included in the definition of polypeptides.
[0042] As used herein, a "fragment" of a sequence refers to a portion of the sequence. For example, a fragment of a nucleic acid sequence refers to a portion of the nucleic acid sequence, and a fragment of an amino acid sequence refers to a portion of the amino acid sequence.
[0043] As used herein, a "variant" of a sequence is a polynucleotide or polypeptide that is respectively different from a reference polynucleotide or polypeptide, but retains the basic characteristics. A typical variant of a polynucleotide differs from another reference polynucleotide in its nucleic acid sequence, and this difference in the nucleic acid sequence may or may not change the amino acid sequence of the polypeptide encoded by the reference polynucleotide. A typical variant of a polypeptide differs from another reference polypeptide in its amino acid sequence. Generally, the differences are limited, so the sequences of the reference polypeptide and the variant are overall very similar and identical in many regions. The variant polypeptide and the reference polypeptide may differ in their amino acid sequences by one or more substitutions, additions, deletions in any combination. The substituted or inserted amino acid residues may or may not be residues encoded by the genetic code. Variants of polynucleotides or polypeptides may be naturally occurring, such as allelic variations, or may be naturally occurring variants that are unknown. Non-naturally occurring polynucleotide and polypeptide variants can be produced by mutagenesis techniques, direct synthesis, and other recombinant methods known to those skilled in the art.
[0044] Amino acids are typically classified by the nature of their side chains. For example, the side chain can make an amino acid a weak acid (such as amino acids D and E) or a weak base (such as amino acids K, R, and H); if the side chain is polar, the amino acid becomes a hydrophilic substance (such as amino acids L and I), or if the side chain is non-polar, the amino acid becomes a hydrophobic substance (such as amino acids S and C).
[0045] As used herein, an "aliphatic amino acid" has a side chain that is an aliphatic group. The aliphatic group causes the amino acid to be non-polar and hydrophobic. The aliphatic group is preferably an unsubstituted branched or linear alkyl group. Non-limiting examples of aliphatic amino acids are A (alanine), V (valine), L (leucine), I (isoleucine), M (methionine), D (aspartic acid), E (glutamic acid), K (lysine), R (arginine), G (glycine), S (serine), T (threonine), C (cysteine), N (asparagine), Q (glutamine).
[0046] As used herein, a "non-polar amino acid" has a non-polar side chain that makes the amino acid hydrophobic. Non-limiting examples of non-polar amino acids are A (alanine), V (valine), L (leucine), I (isoleucine), F (phenylalanine), W (tryptophan), M (methionine), P (proline), and G (glycine).
[0047] As described in the present application, a "polar amino acid" has a polar side chain that renders the amino acid hydrophilic. Non-limiting examples of polar amino acids are T (threonine), S (serine), C (cysteine), N (asparagine), Q (glutamine), Y (tyrosine), K (lysine), R (arginine), H (histidine), D (aspartic acid), and E (glutamic acid). Polar amino acids can be classified into polar uncharged amino acids or polar charged amino acids.
[0048] As described in the present application, a "polar uncharged amino acid" has a polar side chain with an uncharged residue. Non-limiting examples of polar uncharged amino acids are T (threonine), S (serine), C (cysteine), N (asparagine), Q (glutamine), and Y (tyrosine).
[0049] As described in the present application, a "polar charged amino acid" has a polar side chain with at least one charged residue. Non-limiting examples of polar charged amino acids are K (lysine), R (arginine), H (histidine), D (aspartic acid), and E (glutamic acid). Polar charged amino acids can be classified into positively charged amino acids or negatively charged amino acids.
[0050] As described in the present application, a "positively charged amino acid" has a polar side chain with at least one positively charged residue. Non-limiting examples of positively charged amino acids are K (lysine), R (arginine), and H (histidine).
[0051] As described in the present application, a "negatively charged amino acid" has a polar side chain with at least one negatively charged residue. Non-limiting examples of negatively charged amino acids are D (aspartic acid) and E (glutamic acid).
[0052] The term "family" as used in the present application refers to a group of nucleic acids or proteins that are generated from the same ancestor through replication and mutation and have a relatively high structural similarity, and usually have related or even the same functions.
[0053] The term "nuclease" as described in the present application refers to an enzyme that can cleave phosphodiester bonds. Nucleases hydrolyze the phosphodiester bonds in the nucleic acid backbone. The term "endonuclease" as described in the present application refers to an enzyme that can cleave phosphodiester bonds between nucleotides.
[0054] The term "guide RNA" as described in the present application refers to any RNA molecule that can form a complex with the nuclease described in the present application. For example, a guide RNA can be a molecule that recognizes a target gene. In some embodiments of the present application, the guide RNA includes a reRNA and a targeting sequence, wherein the reRNA can bind to a specific nuclease, and the targeting sequence can be designed to be complementary to the target strand of the target gene.
[0055] As used herein, the term "transposon-associated motif" (TAM) refers to a short nucleotide sequence adjacent to a target gene, which can be recognized by the complex formed by the nuclease and guide RNA described in this application. If the target gene is not adjacent to the transposon-associated motif, the nuclease may not be able to successfully recognize the target gene. The sequence and length of the transposon-associated motif in this application may vary depending on the nuclease.
[0056] As used herein, the terms "target gene", "target sequence", "target nucleic acid", "gene of interest", "sequence of interest" and "nucleic acid of interest" are used interchangeably and refer to nucleotide sequences on chromosomal DNA, chloroplast DNA, mitochondrial DNA, plasmid DNA or any other DNA molecule in the cell genome, which can be recognized, bound and selectively cleaved by the complex formed by the nuclease and guide RNA described in this application.
[0057] As used herein, the term "nucleic acid construct" is defined herein as a single-stranded or double-stranded nucleic acid molecule, preferably an artificially constructed nucleic acid molecule. Optionally, the nucleic acid construct further comprises one or more regulatory sequences operably linked thereto, which can direct the expression of the coding sequence in a suitable host cell under its compatible conditions. The term "expression" should be understood to include any steps involved in the production of a protein or polypeptide, including but not limited to transcription, post-transcriptional modification, translation, post-translational modification and secretion. The term "regulatory sequence" includes all components necessary or advantageous for the expression of the polypeptide / protein of this application. Each regulatory sequence may be native or foreign to the nucleic acid sequence encoding the protein or polypeptide. These regulatory sequences include but are not limited to leader sequences, polyadenylation sequences, propeptide sequences, promoters, signal sequences and transcription terminators. At a minimum, the regulatory sequence should include a promoter and transcription and translation termination signals. To introduce specific restriction sites for ligating the regulatory sequence to the coding region of the nucleic acid sequence encoding the protein or polypeptide, the regulatory sequence with a linker can be provided.
[0058] As used herein, the term "promoter" refers to a polynucleotide sequence that can control the transcription of a coding sequence. The promoter sequence includes specific sequences sufficient to enable RNA polymerase to recognize, bind and initiate transcription. In addition, the promoter sequence may include sequences that optionally regulate the recognition, binding and transcriptional initiation activity of RNA polymerase in the nucleic acid construct provided in this application. The promoter can affect the transcription of a gene located on the same nucleic acid molecule as the promoter or a gene located on a different nucleic acid molecule from the promoter.
[0059] As used herein, the term "host cell" includes, but is not limited to, animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells. This term includes the progeny of the original cell into which an exogenous nucleic acid fragment has been introduced. Exemplary host cells include human embryonic kidney cells HEK293T. It should be understood that due to natural, accidental, or intentional mutations, the progeny of a single parental cell may not be identical to the original parent in terms of morphology or in genomic or total DNA complement.
[0060] As used herein, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid molecule to which it is linked. Examples of vectors include, but are not limited to, plasmids, viruses, bacteria, bacteriophages, and insertable DNA fragments. The term "plasmid" refers to a circular double-stranded DNA that can accept an exogenous nucleic acid fragment and can replicate in prokaryotic or eukaryotic cells.
[0061] Nuclease
[0062] The present application provides an isolated nuclease, wherein the nuclease comprises an amino acid sequence represented by the following formula:
[0063] (X1)(X2) a (X3)(X4)(X5) b (X6)(X7) c (X8)(X9) d (X 10 )(X 11 ) e (X 12 )(X 13 ) f (X 14 )(X 15 ) g (X 16 )
[0064] wherein a, b, c, d, e, f, and g are the number of amino acids; (X1), (X3), (X4), (X6), (X8), (X 10 ), (X 12 ), (X 14 ), (X 16 ) are independently polar amino acids or aliphatic amino acids; (X2) is any amino acid, and a is 15 or 16; (X5) is any amino acid, and b is 2; (X7) is any amino acid, and c is 2, 3, or 4; (X9) is any amino acid, and d is 14, 15, 16, 17, or 18; (X 11 ) is any amino acid, and e is 1 or 2; (X 13 ) is any amino acid, and f is 6; and (X 15 ) is any amino acid, and g is 5.
[0065] In some embodiments, (X1) is a positively charged amino acid; (X3) is a polar uncharged amino acid; (X4) is a polar uncharged amino acid; (X6) is a polar uncharged amino acid; (X8) is a polar uncharged amino acid; (X 10 ) is a polar uncharged amino acid; (X 12 ) is a polar uncharged amino acid; (X 14 ) is a negatively charged amino acid; and (X 16 ) is a polar uncharged amino acid. In some embodiments, (X1) is K. In some embodiments, (X3) is S or T. In some embodiments, (X4) is S or T. In some embodiments, (X6) is C. In some embodiments, (X8) is C. In some embodiments, (X 10 ) is C. In some embodiments, (X 12 ) is C. In some embodiments, (X 14 ) is D. In some embodiments, (X 16 ) is N.
[0066] According to embodiments of the present application, an isolated nuclease can be provided, wherein the nuclease has a nuclease sequence selected from the following (i) or a variant sequence of the foregoing nuclease having nuclease activity in (ii)-(iv): (i) at least one of the amino acid sequences shown in any one of SEQ ID NOs: 1-197; (ii) at least one of the sequences obtained by deleting, substituting, inserting or mutating 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids in the amino acid sequence shown in any one of SEQ ID NOs: 1-197; (iii) at least one of the amino acid sequences having at least 70%, 80%, 90%, 95% or 99% identity with the amino acid sequence shown in any one of SEQ ID NOs: 1-197; and (iv) at least one of the sequences obtained by further fusing other sequences to the amino acid sequence shown in any one of SEQ ID NOs: 1-197.
[0067] In some embodiments, the nuclease has a nuclease sequence selected from at least one of the following groups (1)-(9): (1) at least one of the amino acid sequences shown in any one of SEQ ID NO: 52 and 113-147; (2) at least one of the amino acid sequences shown in any one of SEQ ID NO: 27-28, 36-38, 62-85 and 148-171; (3) at least one of the amino acid sequences shown in any one of SEQ ID NO: 13, 86-100 and 105-110; (4) at least one of the amino acid sequences shown in any one of SEQ ID NO: 10-11, 17-19, 29-30 and 174-180; (5) at least one of the amino acid sequences shown in any one of SEQ ID NO: 34, 35, 50, 61 and 181-189; (6) at least one of the amino acid sequences shown in any one of SEQ ID NO: 53 and 190-197; (7) at least one of the amino acid sequences shown in any one of SEQ ID NO: 101, 103, 104 and 112; (8) at least one of the amino acid sequences shown in any one of SEQ ID NO: 7 and 23-25; and (9) at least one of the amino acid sequences shown in any one of SEQ ID NO: 3, 21 and 22.
[0068] In some embodiments, the nuclease has a nuclease sequence selected from at least one of the following groups (1)-(12): (1) at least one of the amino acid sequences shown in any one of SEQ ID NO: 1, 3-4, 6-7, 21-23, 50, 52, 60-61, and 113-147; (2) at least one of the amino acid sequences shown in any one of SEQ ID NO: 14, 27-28, 36-38, 45-48, 59, 62-85, and 148-171; (3) at least one of the amino acid sequences shown in any one of SEQ ID NO: 13, 43, and 86-112; (4) at least one of the amino acid sequences shown in any one of SEQ ID NO: 15-16, 24-25, 32-35, and 181-189; (5) at least one of the amino acid sequences shown in any one of SEQ ID NO: 9, 11, 17-19, 29, and 174-180; (6) at least one of the amino acid sequences shown in any one of SEQ ID NO: 10, 12, 26, 30, 42, and 58; (7) at least one of the amino acid sequences shown in any one of SEQ ID NO: 2, 20, and 31; (8) at least one of the amino acid sequences shown in any one of SEQ ID NO: 8 and 51; (9) at least one of the amino acid sequences shown in any one of SEQ ID NO: 39 and 49; (10) at least one of the amino acid sequences shown in any one of SEQ ID NO: 54 and 55; (11) at least one of the amino acid sequences shown in any one of SEQ ID NO: 53 and 190-197; and (12) at least one of the amino acid sequences shown in any one of SEQ ID NO: 5, 172, and 173.
[0069] In some embodiments, the nuclease belongs to the IS200 / IS605 family. In some embodiments, the nuclease belongs to the IS605, or IS1341 subfamily. In some embodiments, the species source of the nuclease includes Bacteria, or Archaea. In some embodiments, the species source of the nuclease includes Actinobacteria, Aquificae, Bacteroidetes, Candidatus Poribacteria, Chloroflexi, Cyanobacteria, Deinococcus-Thermus, Firmicutes, Planctomycetes, Proteobacteria, Spirochaetes, Tenericutes, Thermotogae, Verrucomicrobia, Candidatus Micrarchaeota, Crenarchaeota, or Euryarchaeota.
[0070] Guide RNA
[0071] According to embodiments of the present application, a guide RNA can be provided, wherein the guide RNA comprises a reRNA, the reRNA comprises a nucleotide sequence shown in any one of SEQ ID NOs: 198-394 or a variant thereof, and the guide RNA is capable of binding to a specific nuclease. In some embodiments, the reRNA comprises at least one nucleotide sequence having at least 70%, 80%, 90%, 95%, or 99% identity to a nucleotide sequence shown in any one of SEQ ID NOs: 198-394. In some embodiments, the reRNA comprises at least one of the nucleotide sequences shown in any one of SEQ ID NOs: 198-394. In some embodiments, the reRNA is at least one of the nucleotide sequences shown in any one of SEQ ID NOs: 198-394.
[0072] In some embodiments, the guide RNA further comprises a targeting sequence, and the targeting sequence can recognize a target gene adjacent to a transposon-related motif. In some embodiments, the length of the targeting sequence is at least one of 10-50, 10-40, 10-30, or 15-25 nucleotides.
[0073] In the present application, the sequence and length of the transposon-related motif may vary depending on the nuclease, and the transposon-related motif can be recognized by the complex formed by the nuclease and the guide RNA described in the present application. In some embodiments, the transposon-related motif comprises a nucleotide sequence represented by the following formula:
[0074] (X 17 ) h (X 18 )(X 19 )A(X 20 )
[0075] wherein, h is the number of nucleotides; A is deoxyadenosine monophosphate; (X 17 ) is any deoxyribonucleotide, and h is 0 or 1; (X 18 ) is deoxycytidine monophosphate or thymidine monophosphate; (X 19 ) is deoxycytidine monophosphate, thymidine monophosphate, or deoxyguanosine monophosphate; and (X 20 ) is any deoxyribonucleotide.
[0076] Nucleic acid, nucleic acid construct
[0077] According to an embodiment of the present application, a nucleic acid can be provided, wherein the nucleic acid encodes the nuclease and / or the guide RNA described in the present application.
[0078] According to an embodiment of the present application, a nucleic acid construct can be provided, which comprises the nucleic acid described in the present application. In some embodiments, the nucleic acid construct further comprises a promoter. The promoter can be any suitable promoter sequence, i.e., a nucleic acid sequence that can be recognized by the host cell expressing the nucleic acid sequence. The promoter sequence contains transcriptional regulatory sequences that mediate the expression of a protein or polypeptide. The promoter can be any nucleic acid sequence that has transcriptional activity in the selected host cell, including mutant, truncated, and hybrid promoters, and can be derived from a gene encoding an extracellular or intracellular protein or polypeptide homologous or heterologous to the host cell. In some embodiments, the promoter includes CMV, EF1a, SV40, PGK, UbC, human β-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.
[0079] In some embodiments, the nucleic acid construct is modified by 5'-end capping and / or 3'-end polyadenylation, and the nucleic acid construct retains the activity of nuclease and / or guide RNA. In some embodiments, the nucleic acid construct is modified by phosphorothioate bond modification, 2'-MOE (2-O-(2-methoxyethyl)), PNA (peptide nucleic acid), GNA (glycerol nucleic acid), LNA (locked nucleic acid), GalNAc (N-acetylgalactosamine), LNP (lipid nanoparticle), PNP (peptide nanoparticle). Methods for modifying nucleic acids are known in the art and are hereby incorporated by reference in their entirety.
[0080] In some embodiments, the nucleic acid construct further comprises a polyA sequence. PolyA tailing signal sequences well-known in the art and various truncated forms of polyA tailing signals can be used in this application.
[0081] In some embodiments, the nucleic acid construct further comprises any transcription termination sequence, i.e., a sequence that can be recognized by the host cell to terminate transcription. The termination sequence is operably linked to the 3'-end of the nucleic acid sequence encoding the protein or polypeptide. Any terminator that can function in the selected host cell can be used in the present invention.
[0082] Optionally, the nucleic acid construct may further comprise a suitable leader sequence, i.e., the untranslated region of mRNA that is important for translation in the host cell. The leader sequence is operably linked to the 5'-end of the nucleic acid sequence encoding the polypeptide. Any leader sequence that can function in the selected host cell can be used in the present invention.
[0083] Optionally, the nucleic acid construct may further comprise a propeptide coding region that encodes an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is called a zymogen or pro-polypeptide. Pro-polypeptides are usually inactive and can be converted into mature active polypeptides by catalytic or autocatalytic cleavage of the propeptide from the pro-polypeptide.
[0084] Optionally, the nucleic acid construct may further comprise regulatory sequences that can regulate the expression of the polypeptide according to the growth of the host cell. Examples of regulatory sequences are those systems that can respond to chemical or physical stimulants (including in the presence of regulatory compounds) to turn on or off gene expression. Other examples of regulatory sequences are those that can cause gene amplification. In these examples, the nucleic acid sequence encoding the protein or polypeptide should be operably linked to the regulatory sequence.
[0085] Composition
[0086] According to an embodiment of the present application, a composition can be provided, wherein the composition comprises: an IS200 / IS605 family nuclease or a functional fragment thereof, or a nucleic acid encoding an IS200 / IS605 family nuclease or a functional fragment thereof, and the nuclease or its functional fragment has endonuclease activity; and a guide RNA, or a nucleic acid encoding the guide RNA, and the guide RNA can bind to a specific nuclease.
[0087] In some embodiments, the composition is selected from at least one of the following groups (1)-(198), and any one of the following groups (1)-(198) comprises: a nuclease-related sequence and a guide RNA-related sequence.
[0088] (1) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:1 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:198.
[0089] (2) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:2 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:199.
[0090] (3) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:3 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:200.
[0091] (4) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:4 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:201.
[0092] (5) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:5 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:202.
[0093] (6) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:6 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:203.
[0094] (7) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:7 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:204;
[0095] (8) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:8 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:205;
[0096] (9) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:9 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:206;
[0097] (10) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:10 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:207;
[0098] (11) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:11 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:208;
[0099] (12) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:12 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:209;
[0100] (13) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:13 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:210;
[0101] (14) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO:14 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO:211;
[0102] (15) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 15 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 212;
[0103] (16) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 16 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 213;
[0104] (17) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 17 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 214;
[0105] (18) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 18 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 215;
[0106] (19) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 19 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 216;
[0107] (20) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 20 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 217;
[0108] (21) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 21 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 218;
[0109] (22) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 22 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 219;
[0110] (23) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 23 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 220;
[0111] (24) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 24 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 221;
[0112] (25) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 25 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 222;
[0113] (26) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 26 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 223;
[0114] (27) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 27 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 224;
[0115] (28) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 28 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 225;
[0116] (29) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 29 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 226;
[0117] (30) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 30 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 227;
[0118] (31) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 31 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 228;
[0119] (32) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 32 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 229;
[0120] (33) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 33 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 230;
[0121] (34) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 34 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 231;
[0122] (35) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 35 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 232;
[0123] (36) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 36 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 233;
[0124] (37) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 37 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 234;
[0125] (38) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 38 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 235;
[0126] (39) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 39 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 236;
[0127] (40) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 40 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 237;
[0128] (41) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 41 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 238;
[0129] (42) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 42 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 239;
[0130] (43) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 43 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 240;
[0131] (44) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 44 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 241;
[0132] (45) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 45 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 242;
[0133] (46) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 46 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 243;
[0134] (47) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 47 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 244;
[0135] (48) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 48 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 245;
[0136] (49) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 49 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 246;
[0137] (50) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 50 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 247;
[0138] (51) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 51 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 248;
[0139] (52) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 52 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 249;
[0140] (53) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 53 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 250;
[0141] (54) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 54 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 251;
[0142] (55) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 55 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 252;
[0143] (56) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 56 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 253;
[0144] (57) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 57 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 254;
[0145] (58) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 58 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 255;
[0146] (59) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 59 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 256;
[0147] (60) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 60 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 257;
[0148] (61) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 61 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 258;
[0149] (62) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 62 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 259;
[0150] (63) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 63 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 260;
[0151] (64) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 64 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 261;
[0152] (65) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 65 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 262;
[0153] (66) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 66 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 263;
[0154] (67) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 67 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 264;
[0155] (68) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 68 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 265;
[0156] (69) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 69 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 266;
[0157] (70) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 70 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 267;
[0158] (71) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 71 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 268;
[0159] (72) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 72 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 269;
[0160] (73) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 73 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 270;
[0161] (74) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 74 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 271;
[0162] (75) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 75 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 272;
[0163] (76) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 76 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 273;
[0164] (77) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 77 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 274;
[0165] (78) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 78 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 275;
[0166] (79) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 79 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 276;
[0167] (80) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 80 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 277;
[0168] (81) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 81 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 278;
[0169] (82) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 82 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 279;
[0170] (83) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 83 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 280;
[0171] (84) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 84 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 281;
[0172] (85) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 85 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 282;
[0173] (86) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 86 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 283;
[0174] (87) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 87 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 284;
[0175] (88) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 88 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 285;
[0176] (89) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 89 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 286;
[0177] (90) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 90 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 287;
[0178] (91) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 91 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 288;
[0179] (92) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 92 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 289;
[0180] (93) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 93 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 290;
[0181] (94) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 94 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 291;
[0182] (95) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 95 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 292;
[0183] (96) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 96 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 293;
[0184] (97) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 97 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 294;
[0185] (98) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 98 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 295;
[0186] (99) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 99 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 296;
[0187] (100) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 100 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 297;
[0188] (101) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 101 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 298;
[0189] (102) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 102 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 299;
[0190] (103) The nuclease-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 103 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 300;
[0191] (104) The nuclease-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 104 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 301;
[0192] (105) The nuclease-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 105 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 302;
[0193] (106) The nuclease-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 106 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 303;
[0194] (107) The nuclease-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 107 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 304;
[0195] (108) The nuclease-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 108 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 305;
[0196] (109) The nuclease-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 109 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 306;
[0197] (110) The nuclease-related sequence is an amino acid sequence containing the sequence shown in SEQ ID NO: 110 or a nucleic acid encoding the amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence containing the sequence shown in SEQ ID NO: 307;
[0198] (111) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 111 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 308;
[0199] (112) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 112 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 309;
[0200] (113) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 113 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 310;
[0201] (114) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 114 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 311;
[0202] (115) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 115 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 312;
[0203] (116) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 116 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 313;
[0204] (117) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 117 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 314;
[0205] (118) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 118 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 315;
[0206] (119) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 119 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 316;
[0207] (120) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 120 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 317;
[0208] (121) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 121 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 318;
[0209] (122) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 122 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 319;
[0210] (123) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 123 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 320;
[0211] (124) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 124 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 321;
[0212] (125) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 125 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 322;
[0213] (126) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 126 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 323;
[0214] (127) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 127 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 324;
[0215] (128) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 128 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 325;
[0216] (129) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 129 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 326;
[0217] (130) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 130 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 327;
[0218] (131) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 131 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 328;
[0219] (132) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 132 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 329;
[0220] (133) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 133 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 330;
[0221] (134) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 134 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 331;
[0222] (135) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 135 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 332;
[0223] (136) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 136 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 333;
[0224] (137) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 137 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 334;
[0225] (138) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 138 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 335;
[0226] (139) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 139 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 336;
[0227] (140) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 140 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 337;
[0228] (141) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 141 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 338;
[0229] (142) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 142 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 339;
[0230] (143) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 143 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 340;
[0231] (144) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 144 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 341;
[0232] (145) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 145 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 342;
[0233] (146) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 146 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 343;
[0234] (147) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 147 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 344;
[0235] (148) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 148 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 345;
[0236] (149) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 149 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 346;
[0237] (150) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 150 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 347;
[0238] (151) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 151 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 348;
[0239] (152) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 152 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 349;
[0240] (153) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 153 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 350;
[0241] (154) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 154 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 351;
[0242] (155) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 155 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 352;
[0243] (156) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 156 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 353;
[0244] (157) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 157 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 354;
[0245] (158) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 158 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 355;
[0246] (159) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 159 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 356;
[0247] (160) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 160 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 357;
[0248] (161) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 161 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 358;
[0249] (162) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 162 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 359;
[0250] (163) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 163 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 360;
[0251] (164) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 164 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 361;
[0252] (165) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 165 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 362;
[0253] (166) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 166 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 363;
[0254] (167) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 167 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 364;
[0255] (168) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 168 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 365;
[0256] (169) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 169 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 366;
[0257] (170) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 170 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 367;
[0258] (171) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 171 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 368;
[0259] (172) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 172 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 369;
[0260] (173) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 173 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 370;
[0261] (174) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 174 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 371;
[0262] (175) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 175 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 372;
[0263] (176) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 176 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 373;
[0264] (177) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 177 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 374;
[0265] (178) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 178 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 375;
[0266] (179) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 179 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 376;
[0267] (180) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 180 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 377;
[0268] (181) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 181 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 378;
[0269] (182) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 182 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 379;
[0270] (183) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 183 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 380;
[0271] (184) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 184 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 381;
[0272] (185) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 185 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 382;
[0273] (186) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 186 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 383;
[0274] (187) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 187 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 384;
[0275] (188) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 188 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 385;
[0276] (189) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 189 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 386;
[0277] (190) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 190 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 387;
[0278] (191) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 191 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 388;
[0279] (192) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 192 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 389;
[0280] (193) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 193 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 390;
[0281] (194) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 194 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 391;
[0282] (195) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 195 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 392;
[0283] (196) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 196 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 393;
[0284] (197) The nuclease-related sequence is an amino acid sequence comprising the sequence shown in SEQ ID NO: 197 or a nucleic acid encoding said amino acid sequence; and the guide RNA-related sequence is a nucleotide sequence comprising the sequence shown in SEQ ID NO: 394;
[0285] (198) A variant of any one of groups (1)-(197) above,
[0286] wherein the nuclease-related sequence is an amino acid sequence of a variant of each group of nucleases or a nucleic acid encoding said variant, and the variant has a variant sequence of the foregoing nuclease having nuclease activity selected from the following (i)-(iii):
[0287] (i) at least one of the sequences obtained by deleting, substituting, inserting, or mutating 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids in the amino acid sequence of each group of nucleases;
[0288] (ii) at least one of the amino acid sequences having at least 70%, 80%, 90%, 95%, or 99% identity with any of the amino acid sequences shown in SEQ ID NO: 1-197; and
[0289] (iii) at least one of the sequences obtained by further fusing other sequences to any of the amino acid sequences shown in SEQ ID NO: 1-197.
[0290] In some embodiments, the guide RNA-related sequence further comprises a targeting sequence that can recognize a target gene adjacent to a transposon-related motif. In some embodiments, the length of the targeting sequence is at least one of 10-50, 10-40, 10-30, or 15-25 nucleotides.
[0291] In the present application, the sequence and length of the transposon-related motif may vary depending on the nuclease, and the transposon-related motif can be recognized by the complex formed by the nuclease and the guide RNA described in the present application. In some embodiments, the transposon-related motif comprises a nucleotide sequence represented by the following formula:
[0292] (X 17 ) h (X 18 )(X 19 )A(X 20 )
[0293] where h is the number of nucleotides; A is deoxyadenosine monophosphate; (X 17 ) is any deoxyribonucleotide, and h is 0 or 1; (X 18 ) is deoxycytidine monophosphate or thymidine monophosphate; (X 19 ) is deoxycytidine monophosphate, thymidine monophosphate, or deoxyguanosine monophosphate; and (X 20 ) is any deoxyribonucleotide.
[0294] The target genes in the present application include any gene of interest, such as a natural functional protein gene, an artificial chimeric gene, or a non-coding RNA gene. In some embodiments, the natural functional protein gene includes a luciferase reporter gene, a luciferase gene, and a resistance gene. In some embodiments, the artificial chimeric gene includes a chimeric antigen receptor gene. In some embodiments, the luciferase reporter gene includes a gene encoding green fluorescent protein, red fluorescent protein, blue fluorescent protein, or yellow fluorescent protein. In some embodiments, the luciferase gene includes a gene encoding firefly luciferase or Renilla luciferase. In some embodiments, the resistance gene includes a gene encoding puromycin resistance, G418 resistance, kanamycin resistance, tetracycline resistance, or bleomycin resistance.
[0295] In some embodiments, the nucleic acid encoding an amino acid sequence and / or the nucleic acid encoding a guide RNA further comprises a promoter. The promoter can be any suitable promoter sequence, i.e., a nucleic acid sequence recognizable by the host cell that expresses the nucleic acid sequence. The promoter sequence contains transcriptional regulatory sequences that mediate the expression of a protein or polypeptide. The promoter can be any nucleic acid sequence that is transcriptionally active in the selected host cell, including mutant, truncated, and chimeric promoters, and can be derived from a gene encoding an extracellular or intracellular protein or polypeptide that is homologous or heterologous to the host cell. In some embodiments, the promoter includes CMV, EF1a, SV40, PGK, UbC, human β-actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL. In some embodiments, the nucleic acid encoding an amino acid sequence and / or the nucleic acid encoding a guide RNA further comprises a polyA sequence. Well-known polyA tailing signal sequences in the art and various truncated forms of polyA tailing signal sequences can be used in the present application.
[0296] In some embodiments, the nucleic acid encoding an amino acid sequence and / or the nucleic acid encoding a guide RNA further comprises any transcriptional termination sequence to control the expression of the foreign nucleic acid fragment, i.e., a sequence that can be recognized by the host cell to terminate transcription. Any terminator that can function in the selected host cell can be used in the present invention.
[0297] In some embodiments, the nucleic acid encoding an amino acid sequence and / or the nucleic acid encoding a guide RNA further comprises any transcriptional termination sequence, i.e., a sequence that can be recognized by the host cell to terminate transcription. The termination sequence is operably linked to the 3' end of the nucleic acid sequence encoding the protein or polypeptide. Any terminator that can function in the selected host cell can be used in the present invention.
[0298] Optionally, the nucleic acid encoding the amino acid sequence and / or the nucleic acid encoding the guide RNA may further comprise a suitable leader sequence, i.e., the untranslated region of the mRNA that is important for translation in the host cell. The leader sequence is operably linked to the 5'-end of the nucleic acid sequence encoding the polypeptide. Any leader sequence that can function in the selected host cell can be used in the present invention.
[0299] Optionally, the nucleic acid encoding the amino acid sequence and / or the nucleic acid encoding the guide RNA may further comprise a propeptide coding region that encodes an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is called a zymogen or pro-polypeptide. Pro-polypeptides are usually inactive and can be converted into mature active polypeptides by catalytic or autocatalytic cleavage of the propeptide from the pro-polypeptide.
[0300] Optionally, the nucleic acid encoding the amino acid sequence and / or the nucleic acid encoding the guide RNA may further comprise regulatory sequences that can regulate the expression of the polypeptide according to the growth of the host cell. Examples of regulatory sequences are those systems that can respond to chemical or physical stimuli (including in the presence of regulatory compounds) to turn on or off gene expression. Other examples of regulatory sequences are those that can cause gene amplification. In these examples, the nucleic acid sequence encoding the protein or polypeptide should be operably linked to the regulatory sequence.
[0301] Recombinant vector, recombinant host cell and kit
[0302] According to an embodiment of the present application, a recombinant vector can be provided, wherein the recombinant vector comprises a nucleic acid encoding the nuclease described in the present application, the guide RNA described in the present application, the nucleic acid described in the present application, the nucleic acid construct described in the present application, or the composition described in the present application. The recombinant vector can be any suitable vector. In some embodiments, the recombinant vector includes but is not limited to a recombinant cloning vector, a recombinant eukaryotic expression plasmid, or a recombinant viral vector. In some embodiments, the recombinant eukaryotic expression plasmid includes pcDNA3.1, pCMV, pUC18, pUC19, pUC57, pBAD, pET, pENTR, pGenlenti, or pAAV. In some embodiments, the recombinant viral vector includes a recombinant adenovirus vector, a recombinant adeno-associated virus vector, a recombinant retrovirus vector, a recombinant herpes simplex virus vector, or a recombinant vaccinia virus vector. The recombinant vectors of the present invention can be constructed by methods well known in the art. For example, appropriate restriction sites can be added to both ends of the nucleic acid construct of the present invention according to the restriction sites contained in the backbone vector used, and then loaded into the backbone vector.
[0303] According to embodiments of the present application, a recombinant host cell can be provided, wherein the recombinant host cell contains the nuclease, the guide RNA, the nucleic acid, the nucleic acid construct, the composition, or the recombinant vector described in the present application. The recombinant host cell can be any host cell to which the nuclease can be applied. In some embodiments, the recombinant host cell includes, but is not limited to, animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells. In some embodiments, the animal cells include mammalian cells. In some embodiments, the mammalian cells include primary cells (such as mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells), immortalized cell lines (such as HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial / epithelial / fibroblast / keratinocyte / ductal / cell lines), cancer cell lines (such as Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, Renca), embryonic stem cell lines (such as H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, D3) and their differentiated cells, or induced pluripotent stem cell lines and their differentiated cells. In some embodiments, the plant cells include monocotyledonous plant cells or dicotyledonous plant cells. In some embodiments, the monocotyledonous plant cells or dicotyledonous plant cells include rice cells, corn cells, or soybean cells.
[0304] According to embodiments of the present application, a kit can be provided, wherein the kit contains the nuclease, the guide RNA, the nucleic acid, the nucleic acid construct, the composition, the recombinant vector, or the recombinant host cell described in the present application.
[0305] Method and use
[0306] The nuclease-based gene editing tools and methods provided by the present application can be applied to multiple fields such as gene therapy, molecular breeding of animals and plants, transformation of industrial microorganisms, transformation of model animals, and scientific research. Especially in the field of gene therapy, it can be applied to gene knockout based on DNA double-strand breaks in the human genome.
[0307] According to an embodiment of the present application, a method for introducing a double-strand break into a target gene of a host cell can be provided, wherein the method comprises: delivering the nuclease, the guide RNA, the nucleic acid, the nucleic acid construct, the composition, or the recombinant vector described in the present application into the host cell.
[0308] According to an embodiment of the present application, a method for deleting, replacing, or inserting a target gene of a host cell can be provided, wherein the method comprises: delivering the nuclease, the guide RNA, the nucleic acid, the nucleic acid construct, the composition, or the recombinant vector described in the present application into the host cell.
[0309] According to an embodiment of the present application, a method for obtaining a host cell with a target gene deleted, replaced, or inserted can be provided, wherein the method comprises: delivering the nuclease, the guide RNA, the nucleic acid, the nucleic acid construct, the composition, or the recombinant vector described in the present application into the host cell.
[0310] The method of delivering into the host cell can be any suitable method. In some embodiments, the delivery methods include but are not limited to cationic liposome delivery, lipid nanoparticle delivery, cationic polymer delivery, vesicle-exosome delivery, gold nanoparticle delivery, polypeptide and protein delivery, retroviral delivery, lentiviral delivery, adenoviral delivery, adeno-associated viral delivery, electroporation, Agrobacterium infection, or gene gun. The methods of cell transfection and culture are conventional methods in the art, and suitable transfection and culture methods can be selected according to different cell types.
[0311] The host cell can be any host cell to which the nuclease can be applied. In some embodiments, the host cell includes, but is not limited to, animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells. In some embodiments, the animal cells include mammalian cells. In some embodiments, the mammalian cells include primary cells (e.g., mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells), immortalized cell lines (e.g., HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial / epithelial / fibroblast / keratinocyte / ductal cell lines), cancer cell lines (e.g., Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, Renca), embryonic stem cell lines (e.g., H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, D3) and their differentiated cells, or induced pluripotent stem cell lines and their differentiated cells. In some embodiments, the plant cells include monocotyledonous plant cells or dicotyledonous plant cells. In some embodiments, the monocotyledonous plant cells or dicotyledonous plant cells include rice cells, corn cells, or soybean cells.
[0312] According to an embodiment of the present application, there can be provided the use of the nuclease, the guide RNA, the nucleic acid, the nucleic acid construct, the composition, the recombinant vector, or the recombinant host cell described in the present application in introducing a double-strand break into a target gene of a host cell.
[0313] According to an embodiment of the present application, there can be provided the use of the nuclease, the guide RNA, the nucleic acid, the nucleic acid construct, the composition, the recombinant vector, or the recombinant host cell described in the present application in deleting, replacing, or inserting a target gene of a host cell.
[0314] The host cell can be any host cell to which the nuclease can be applied. In some embodiments, the host cell includes, but is not limited to, animal cells, plant cells, algal cells, fungal cells, yeast cells, or bacterial cells. In some embodiments, the animal cells include mammalian cells. In some embodiments, the mammalian cells include primary cells (such as mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells), immortalized cell lines (such as HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial / epithelial / fibroblast / keratinocyte / ductal / cell lines), cancer cell lines (such as Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, Renca), embryonic stem cell lines (such as H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, D3) and their differentiated cells, or induced pluripotent stem cell lines and their differentiated cells. In some embodiments, the plant cells include monocotyledonous plant cells or dicotyledonous plant cells. In some embodiments, the monocotyledonous plant cells or dicotyledonous plant cells include rice cells, corn cells, or soybean cells.
[0315] According to the embodiments of the present application, there can be provided the use of the nuclease, guide RNA, nucleic acid, nucleic acid construct, composition, recombinant vector, or recombinant host cell described in the present application in the preparation of drugs or formulations for gene therapy, cell therapy, genomic research, stem cell induction, and post-induction differentiation.
[0316] The various embodiments and preferred options described above for the present application can be combined with each other (as long as they are not inherently contradictory to each other), and are applicable to the uses of the present application. All the various embodiments formed by such combinations are regarded as part of the present application.
[0317] Example
[0318] The following describes exemplary embodiments of the present application in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding. It should be understood that they are considered to be merely exemplary and are in no way intended to limit the scope of protection of the present application. The scope of protection of the present application is defined only by the claims. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present application. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.
[0319] Unless otherwise specified, the reagents and instruments used in the following examples are all traditional products that can be obtained commercially. Unless otherwise specified, the experiments are carried out under traditional conditions or the conditions recommended by the manufacturer.
[0320] Example 1: Construction of nuclease activity detection system
[0321] We established a set of RGS dual-fluorescence surrogate reporter systems to verify the activity of candidate nucleases.
[0322] Plasmid 1 consists of a complete set of elements capable of transcribing and expressing candidate nuclease proteins, including a constitutive promoter CMV (the sequence is shown in SEQ ID NO: 405) that can initiate transcription in eukaryotic cells, the candidate nuclease sequence (shown in Table 1), the 5'-nuclear localization signal peptide sequence (the sequence is shown in SEQ ID NO: 406), the 3'-nuclear localization signal peptide sequence (the sequence is shown in SEQ ID NO: 407), the polyA sequence that terminates transcription (the sequence is shown in SEQ ID NO: 408), and the ampicillin resistance gene sequence (the sequence is shown in SEQ ID NO: 409).
[0323] Construction method of plasmid 1: The amino acid sequence (or nucleotide sequence) of the candidate nuclease protein was submitted to BGI Tech Solutions (Beijing Liuhe) Co., Ltd. for traditional gene synthesis. An EcoRI cleavage site was inserted at the 5' end of the upstream sequence, and a BamH1 cleavage site was inserted at the 3' end of the downstream sequence. The plasmid construction was also responsible by the gene synthesis company. The specific construction method was as follows: 1. Preparation of the vector. The plasmid backbone of the pcDNA3.1 plasmid vector was digested with two restriction endonucleases, EcoRI and BamHI, at their single cleavage sites. The linearized plasmid vector fragment was obtained by agarose gel electrophoresis and the enzyme digestion band was cut from the gel for recovery to obtain a purified linearized plasmid vector fragment. 2. Ligation. The nucleotide sequence of the candidate nuclease protein obtained by traditional gene synthesis was ligated to the linearized pcDNA3.1 vector fragment using T4 DNA ligase. 3. Transformation and verification. Monoclonal transformants were obtained through an ampicillin-resistant LB screening agar plate, and the correct clones were confirmed by sequencing and used as candidate plasmids.
[0324] Plasmid 2 contains the reRNA sequence (as shown in Table 1). A 20-nt targeting sequence, GCTCGGAGATCATCATTGCG, the U6 promoter (the sequence is as shown in SEQ ID NO: 410), the PBR322 origin of replication (the sequence is as shown in SEQ ID NO: 411), and the ampicillin resistance gene sequence (the sequence is as shown in SEQ ID NO: 409) were inserted at the 3' end of the reRNA sequence.
[0325] Construction method of plasmid 2: The guide reRNA was submitted to Beijing Tsingke Biotech Co., Ltd. or General Biosystems (Anhui) Co., Ltd. for traditional gene synthesis. The plasmid construction was also responsible by the gene synthesis company. The specific construction method was as follows: 1. Preparation of the vector. The pUC19-U6 vector was digested with BbsI. The linearized plasmid vector fragment was obtained by agarose gel electrophoresis and the enzyme digestion band was cut from the gel for recovery to obtain a purified linearized plasmid vector fragment. 2. Ligation. The nucleotide sequence of the guide reRNA obtained by gene synthesis was ligated to the linearized pUC19-U6 vector fragment using seamless cloning. 3. Transformation and verification. Monoclonal transformants were obtained through an ampicillin-resistant LB screening agar plate, and the correct clones were confirmed by sequencing and used as candidate plasmids.
[0326] Plasmid 3 contains the TAM sequence (shown in Table 1), with a 20-nt targeting sequence GCTCGGAGATCATCATTGCG inserted at the 3' end of the TAM sequence, the CMV promoter (the sequence is as shown in SEQ ID NO: 412), the ampicillin resistance gene sequence (the sequence is as shown in SEQ ID NO: 409), and an alternative reporter gene. The alternative reporter gene can encode two fluorescent proteins (the RFP sequence shown in SEQ ID NO: 413 and the GFP sequence shown in SEQ ID NO: 414). By inserting an endonuclease downstream of RFP and upstream of GFP, the TAM and the 20-nt targeting sequence at the 3' end of TAM can be recognized. When there is no endonuclease activity in the detection system, only RFP is expressed by the reporter gene to indicate the reference gene expression level of the reporter system, while GFP is designed outside the open reading frame (ORF) and thus not expressed. When the candidate has endonuclease activity, it can induce a double-strand break at the target site in front of GFP, resulting in a frameshift mutation in the reading frame when the DNA is repaired by non-homologous end joining (NHEJ), causing GFP to shift from an out-of-frame state to an in-frame state and start to be expressed. The stronger the cleavage activity of the nuclease, the higher the proportion of GFP expressed after frameshift. Therefore, the expression intensity of GFP is positively correlated with the cleavage activity of the nuclease. The working mode of the detection system is as Figure 1 shown.
[0327] Construction method of plasmid 3: The TAM, 20-nt targeting sequence, and the 5' end of the upstream sequence were inserted into the ECoRI cleavage site, and the 3' end of the downstream sequence was inserted into the BamH1 cleavage site by oligonucleotide synthesis for overall synthesis. The specific construction steps are as follows: 1. Preparation of the vector. The plasmid backbone of the RGS-pcDNA3.1 plasmid vector was digested with two restriction endonucleases, ECoRI and BamHI, at their single restriction endonuclease cleavage sites. The linearized plasmid vector fragment was obtained by agarose gel electrophoresis and cut from the gel for recovery to obtain a purified linearized plasmid vector fragment. 2. Ligation. The guiding reRNA nucleotide sequence synthesized by gene synthesis was ligated to the linearized pUC19-U6 vector fragment by seamless cloning ligation. 3. Transformation and verification. Monoclonal transformants were obtained through an ampicillin-resistant LB screening agar plate, and the correct clones were confirmed by sequencing and used as candidate plasmids.
[0328] Table 1 Sequences related to plasmid construction
[0329]
[0330]
[0331]
[0332]
[0333]
[0334]
[0335]
[0336] Example 2: Nuclease activity detection
[0337] 2.1 Cell treatment:
[0338] After culturing HEK293T cells (commercially purchased) to the logarithmic growth phase, they were trypsinized into single cells with 0.25% trypsin (Thermo Fisher Scientific), and added to 96-well cell culture plates pre-coated with PDL (Sigma-Aldrich) at a cell concentration of 3×10 4 cells / well, and cultured overnight at 5% CO2 and 37°C.
[0339] 2.2 Cell transfection:
[0340] The three functional plasmids (nuclease plasmid, reRNA-targeting sequence plasmid, RGS dual-fluorescent reporter system plasmid) described in Example 1 were co-transfected into HEK293T cells. Among them, 60 ng of nuclease plasmid, 40 ng of reRNA-targeting sequence plasmid, and 100 ng of RGS dual-fluorescent reporter system plasmid were added to the 96-well cell culture plates respectively, and lipofectamine TM 2000 (Invitrogen, catalog number 11668019) was used for transfection at a ratio of transfection reagent volume (μL): plasmid mass (μg) of 2:1.
[0341] 2.3 Obtaining results
[0342] After transfection, the cells were cultured for 48 h, then trypsinized and collected, and detected by flow cytometry. The final screening results were analyzed based on the GFP positive expression level.
[0343] 2.4 Detection results
[0344] The results of nuclease activity were obtained by flow cytometry, and the data are as Figure 2 shown. The vertical axis in the figure represents the RFP expression level (%), and the horizontal axis represents the GFP expression level (%), which reflects the cleavage activity of the nuclease. In addition, the statistical results of the GFP expression level reflecting all nuclease activities are as Figure 3as shown in Table 2. The results show that 197 nucleases of the present application (TP_A_1, TP_A_2, TP_A_8, TP_A_12, TP_A_18, TP_B_18, TP_B_41, TP_B_46, TP_B_70, TP_B_71, TP_B_72, TP_B_73, TP_C_23, TP_C_67, TP_C_70, TP_C_74, TP_D_1, TP_D_3, TP_D_4, TP_D_8, TP_D_17, TP_D_18, TP_D_23, TP_D_24, TP_D_25, TP_D_27, TP_D_30, TP_D_32, TP_D_40, TP_D_43, TP_D_51, TP_D_59, TP_D_61, TP_D_66, TP_D_67, TP_D_71, TP_D_72, TP_D_73, TP_E_2, TP_E_15, TP_E_17, TP_E_48, TP_F_56, TP_F_71, TP_F_77, TP_F_80, TP_F_83, TP_F_85, TP_G_14, TP_G_19, TP_G_20, TP_G_24, TP_G_43, TP_G_52, TP_G_53, TP_G_61, TP_G_66, TP_G_72, TP_G_75, TP_G_83, TP_G_84, TP_H_1, TP_H_3, TP_H_4, TP_H_5, TP_H_6, TP_H_9, TP_H_11, TP_H_12, TP_H_13, TP_H_15, TP_H_18, TP_H_19, TP_H_20, TP_H_21, TP_H_23, TP_H_24, TP_H_30, TP_H_31, TP_H_32, TP_H_34, TP_H_38, TP_H_39, TP_H_40, TP_H_43, TP_I_1, TP_I_2, TP_I_3, TP_I_4, TP_I_5, TP_I_6, TP_I_7, TP_I_8, TP_I_9, TP_I_10, TP_I_11, TP_I_12, TP_I_13, TP_I_15, TP_I_16, TP_I_17, TP_I_18, TP_I_19, TP_I_20, TP_I_21, TP_I_22, TP_I_24, TP_I_25, TP_I_26, TP_I_29, TP_I_31, TP_I_35, TP_I_37, TP_I_38, TP_I_40, TP_I_41, TP_I_44, TP_I_45, TP_I_46, TP_I_47, TP_I_48, TP_I_49, TP_I_50, TP_I_51, TP_I_52,TP_I_53, TP_I_55, TP_I_56, TP_I_58, TP_I_59, TP_I_61, TP_I_62, TP_I_64, TP_I_65, TP_I_66, TP_I_67, TP_I_70, TP_I_71, TP_I_76, TP_I_77, TP_I_79, TP_I_80, TP_I_82, TP_I_84, TP_I_85, TP_I_86, TP_I_87, TP_L_1, TP_L_4, TP_L_5, TP_L_8, TP_L_9, TP_L_10, TP_L_11, TP_L_12, TP_L_15, TP_L_16, TP_L_17, TP_L_21, TP_L_22, TP_L_24, TP_L_25, TP_L_26, TP_L_27, TP_L_28, TP_L_31, TP_L_32, TP_L_34, TP_L_36, TP_L_37, TP_L_39, TP_M_1, TP_M_3, TP_M_7, TP_M_11, TP_M_14, TP_M_17, TP_M_19, TP_M_20, TP_M_24, TP_M_31, TP_M_32, TP_M_33, TP_M_34, TP_M_35, TP_M_37, TP_M_40, TP_M_41, TP_M_43, TP_M_46, TP_M_49, TP_M_58, TP_M_65, TP_M_66, TP_M_67, TP_M_70 and TP_M_78) have good activity.
[0345] Meanwhile, a large number of inactive or low-cutting-activity nucleases were also found during the screening process (for example, TP_A_24, TP_A_54, TP_B_23, TP_D_44, and TP_F_76 in Table 1 of this application form). Compared with these inactive or low-activity nucleases, the cutting activity of the 197 nucleases of this application is significantly higher.
[0346] In addition, Figure 11 shows the phylogenetic tree of the nucleases in this application based on the protein sequence. The results show that these nucleases cover different branches of the superfamily, and ISDra2 is also included therein.
[0347] Results of nuclease activity in Example 2 of Table 2
[0348]
[0349]
[0350]
[0351]
[0352]
[0353]
[0354]
[0355] Example 3: Detection of editing efficiency at endogenous locus
[0356] 3.1 Construction of plasmids:
[0357] The nuclease plasmid (plasmid 1) contains a complete set of elements capable of transcribing and expressing the candidate nuclease protein, including the constitutive promoter CMV (the sequence is shown in SEQ ID NO: 405) that can initiate transcription in eukaryotic cells, the candidate nuclease sequence (shown in Table 1), the 5'-nuclear localization signal peptide sequence (the sequence is shown in SEQ ID NO: 406), the 3'-nuclear localization signal peptide sequence (the sequence is shown in SEQ ID NO: 407), the polyA sequence that terminates transcription (the sequence is shown in SEQ ID NO: 408), and the ampicillin resistance gene sequence (the sequence is shown in SEQ ID NO: 409). The construction method of plasmid 1 is described in Example 1.
[0358] The reRNA-targeting sequence plasmid (plasmid 4) contains the reRNA sequence (shown in Table 1), inserts a 20nt endogenous gene targeting sequence (shown in Table 3) at the 3' end of the reRNA sequence, the U6 promoter (the sequence is shown in SEQ ID NO: 410), the PBR322 replication origin (the sequence is shown in SEQ ID NO: 411), and the ampicillin resistance gene sequence (the sequence is shown in SEQ ID NO: 409). In addition, different targeting sequences of the endogenous gene (shown in Table 2) can identify different target genes adjacent to the TAM sequence.
[0359] Construction method of plasmid 4:
[0360] 1. Preparation of the vector. The reRNA plasmid (pUC19-U6-reRNA-BbsI_BbsI) containing the BBSI-BBSI fragment was digested with BbsI, and the linearized plasmid vector fragment was obtained by agarose gel electrophoresis and cut from the gel for recovery to obtain the purified linearized plasmid vector fragment. 2. Preparation of the 20nt endogenous gene targeting sequence. First, search for the 20bp DNA sequence adjacent to the 3' end of the TAM sequence in the endogenous gene sequence, and then synthesize the oligonucleotide with the targeting sequence having the BbsI excision end by primer synthesis. Finally, anneal and bond to synthesize the double-stranded oligonucleotide with sticky ends. 3. Ligation. The targeting sequence of the endogenous gene was ligated to the linearized pUC19-U6-reRNA-BbsI_BbsI vector fragment using T4 ligase. 4. Transformation and verification. Monoclonal transformants were obtained by screening on an ampicillin-resistant LB agar plate, and the correct clones were confirmed by sequencing and used as candidate plasmids.
[0361] 3.2 Cell treatment:
[0362] After culturing HEK293T cells (commercially purchased) to the logarithmic growth phase, they were trypsinized into single cells with 0.25% trypsin (Thermo Fisher Scientific), and added to a 48-well cell culture plate pre-coated with PDL (Sigma-Aldrich) at a cell concentration of 1×10 5 cells / well, and cultured overnight at 5% CO2 and 37°C.
[0363] 3.3 Cell transfection:
[0364] The two functional plasmids (nuclease plasmid and reRNA-targeting sequence plasmid) described in 3.1 were co-transfected into HEK293T cells. Among them, 300 ng of the nuclease plasmid and 200 ng of the reRNA-targeting sequence plasmid were added to the 48-well cell culture plate respectively, and lipofectamine TM 2000 (Invitrogen, catalog number 11668019) was used for transfection at a ratio of transfection reagent volume (μL): plasmid mass (μg) of 2:1.
[0365] 3.4 PCR amplification and NGS second-generation sequencing
[0366] After transfection, the cells were cultured for 48 h, then trypsinized and collected, and genomic DNA was extracted. PCR primers were designed close to the targeting sequence of the endogenous gene to amplify a PCR product of about 200 bp in length including the 20nt targeting sequence. The PCR product was sequenced by next-generation sequencing.
[0367] 3.5 Detection results
[0368] The results of endogenous gene editing efficiency are as Figure 9 shown in Table 3.
[0369] By analyzing the sequence data generated by next-generation sequencing technology, the endogenous gene editing activity of the nuclease was determined by counting the base insertions and deletions (Indel%) generated on the target sequence of the endogenous gene. The results showed that the nuclease in this application showed good editing efficiency on different endogenous genes.
[0370] Table 3 Results of endogenous gene editing activity in Example 3
[0371]
[0372]
[0373]
[0374]
[0375]
[0376]
[0377] Example 4: Detection of nuclease activity in rice protoplasts
[0378] In this example, a pair of synthetic YFP gene reporter vectors (Plasmids 5 and 6) were used to evaluate the nuclease activity in rice protoplasts. These vectors were constructed using the method described in Example 1 (as Figure 13 shown).
[0379] Plasmid 5 contains the promoter ZmUBI (SEQ ID NO:593), the candidate nuclease sequence (shown in Table 1), the NOS terminator (SEQ ID NO:594), the promoter OsU6 (SEQ ID NO:595), the reRNA sequence corresponding to the specific nuclease (shown in Table 1), the spacer sequence (SEQ ID NO:596), and the terminator (SEQ ID NO:597).
[0380] In Plasmid 6, the YFP sequence (SEQ ID NO:598) is segmented by the spacer sequence (SEQ ID NO:596) and the TAM sequence (shown in Table 1) (corresponding to the specific nuclease in Plasmid 5). The YFP sequence in the first half overlaps with the YFP sequence in the second half. Plasmid 6 also contains the promoter 35S (SEQ ID NO:599) and the terminator (SEQ ID NO:600).
[0381] After co-transforming plasmid 5 and plasmid 6 into rice protoplasts, once the spacer sequence in plasmid 6 is cleaved by a nuclease, the partially overlapping fragments (derived from the middle region of YFP) promote DSB repair through the homology-dependent DNA repair pathway, thereby restoring the normal YFP gene (as Figure 14 shown). Therefore, the cleavage activity of the nuclease can be evaluated by observing the number of YFP-positive cells.
[0382] The results of YFP fluorescence are as Figure 15 - 31As shown, 197 nucleases of the present application are shown (TP_A_1, TP_A_2, TP_A_8, TP_A_12, TP_A_18, TP_B_18, TP_B_41, TP_B_46, TP_B_70, TP_B_71, TP_B_72, TP_B_73, TP_C_23, TP_C_67, TP_C_70, TP_C_74, TP_D_1, TP_D_3, TP_D_4, TP_D_8, TP_D_17, TP_D_18, TP_D_23, TP_D_24, TP_D_25, TP_D_27, TP_D_30, TP_D_32, TP_D_40, TP_D_43, TP_D_51, TP_D_59, TP_D_61, TP_D_66, TP_D_67, TP_D_71, TP_D_72, TP_D_73, TP_E_2, TP_E_15, TP_E_17, TP_E_48, TP_F_56, TP_F_71, TP_F_77, TP_F_80, TP_F_83, TP_F_85, TP_G_14, TP_G_19, TP_G_20, TP_G_24, TP_G_43, TP_G_52, TP_G_53, TP_G_61, TP_G_66, TP_G_72, TP_G_75, TP_G_83, TP_G_84, TP_H_1, TP_H_3, TP_H_4, TP_H_5, TP_H_6, TP_H_9, TP_H_11, TP_H_12, TP_H_13, TP_H_15, TP_H_18, TP_H_19, TP_H_20, TP_H_21, TP_H_23, TP_H_24, TP_H_30, TP_H_31, TP_H_32, TP_H_34, TP_H_38, TP_H_39, TP_H_40, TP_H_43, TP_I_1, TP_I_2, TP_I_3, TP_I_4, TP_I_5, TP_I_6, TP_I_7, TP_I_8, TP_I_9, TP_I_10, TP_I_11, TP_I_12, TP_I_13, TP_I_15, TP_I_16, TP_I_17, TP_I_18, TP_I_19, TP_I_20, TP_I_21, TP_I_22, TP_I_24, TP_I_25, TP_I_26, TP_I_29, TP_I_31, TP_I_35, TP_I_37, TP_I_38, TP_I_40, TP_I_41, TP_I_44, TP_I_45, TP_I_46, TP_I_47, TP_I_48, TP_I_49, TP_I_50, TP_I_51, TP_I_52,TP_I_53, TP_I_55, TP_I_56, TP_I_58, TP_I_59, TP_I_61, TP_I_62, TP_I_64, TP_I_65, TP_I_66, TP_I_67, TP_I_70, TP_I_71, TP_I_76, TP_I_77, TP_I_79, TP_I_80, TP_I_82, TP_I_84, TP_I_85, TP_I_86, TP_I_87, TP_L_1, TP_L_4, TP_L_5, TP_L_8, TP_L_9, TP_L_10, TP_L_11, TP_L_12, TP_L_15, TP_L_16, TP_L_17, TP_L_21, TP_L_22, TP_L_24, TP_L_25, TP_L_26, TP_L_27, TP_L_28, TP_L_31, TP_L_32, TP_L_34, TP_L_36, TP_L_37, TP_L_39, TP_M_1, TP_M_3, TP_M_7, TP_M_11, TP_M_14, TP_M_17, TP_M_19, TP_M_20, TP_M_24, TP_M_31, TP_M_32, TP_M_33, TP_M_34, TP_M_35, TP_M_37, TP_M_40, TP_M_41, TP_M_43, TP_M_46, TP_M_49, TP_M_58, TP_M_65, TP_M_66, TP_M_67, TP_M_70 and TP_M_78) also have good cleavage activity in rice protoplasts.
[0383] It should be noted that the above are only preferred examples of this application and are not intended to limit this application. Various modifications and changes can be made to this application by those of ordinary skill in the art. Although specific embodiments have been described, for the applicant or other persons skilled in the art, there may exist or currently be unforeseen alternatives, modifications, changes, improvements, and substantial equivalents of the above embodiments. Therefore, the appended claims submitted and the claims that may be modified are intended to cover all such alternatives, modifications, changes, improvements, and substantial equivalents. Importantly, as technology evolves, many of the elements described herein can be replaced by equivalent elements that emerge after this application.
Claims
1. An isolated nuclease, wherein the amino acid sequence of the nuclease is shown in SEQ ID NO:
68.
2. The nuclease according to claim 1, wherein the nuclease belongs to the IS200 / IS605 family.
3. The nuclease according to claim 2, wherein the nuclease belongs to the IS605 or IS1341 subfamily. The nuclease according to claim 1 , wherein the species of origin of the nuclease comprises bacteria. The nuclease according to claim 4 , wherein the species source of the nuclease comprises Firmicutes.
6. A reRNA, wherein the reRNA is the nucleotide sequence shown in SEQ ID NO:
265.
7. A guide RNA, wherein the guide RNA comprises a reRNA, the reRNA has a nucleotide sequence as shown in SEQ ID NO: 265, and the guide RNA is capable of binding to a specific nuclease; The guide RNA also comprises a targeting sequence capable of recognizing a target gene adjacent to the transposon-associated motif.
8. The guide RNA of claim 7, wherein the targeting sequence is 15-25 nucleotides in length.
9. The guide RNA according to claim 7, wherein the transposon-related motif is a nucleotide sequence shown in the following formula: TTTAA.
10. A nucleic acid, wherein the nucleic acid encodes the nuclease according to any one of claims 1-5, the reRNA according to claim 6 and / or the guide RNA according to any one of claims 7-9. A nucleic acid construct comprising the nucleic acid according to claim 10 .
12. A nucleic acid construct according to claim 11, wherein the nucleic acid construct comprises a promoter, wherein the promoter is selected from CMV, EF1a, SV40, PGK, UbC, human beta actin, CAG, TRE, UAS, Ac5, GFAP, polyhedrin promoter, TBG, ALB, ApoEHCR-hAAT, CaMKIIa, GAL1, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, U6, T7, T7lac, Sp6, araBAD, trp, lac, Ptac, or pL.
13. The nucleic acid construct of claim 11, wherein the nucleic acid construct is modified by 5' end capping and / or 3' end polyadenylation, and the nucleic acid construct retains the activity of a nuclease and / or a guide RNA.
14. The nucleic acid construct of claim 11, wherein the nucleic acid construct is modified by phosphorothioate bond modification, 2'-MOE (2-O-(2-methoxyethyl)), PNA (peptide nucleic acid), GNA (glycerol nucleic acid), LNA (locked nucleic acid), GalNAc (N-acetylgalactosamine), LNP (lipid nanoparticle), PNP (peptide nanoparticle).
15. A composition, wherein the composition comprises: IS200 / IS605 family nuclease, or comprising a nucleic acid encoding an IS200 / IS605 family nuclease, wherein the nuclease has endonuclease activity; and A guide RNA, or a nucleic acid comprising a guide RNA encoding the guide RNA, wherein the guide RNA is capable of binding to a specific nuclease; The composition comprises: a nuclease-associated sequence and a guide RNA-associated sequence, wherein the nuclease-associated sequence is an amino acid sequence shown in SEQ ID NO: 68 or a nucleic acid sequence encoding the amino acid sequence; the guide RNA-associated sequence comprises reRNA, and the reRNA is a nucleotide sequence shown in SEQ ID NO: 265; and the guide RNA-associated sequence further comprises a targeting sequence, and the targeting sequence can recognize a target gene adjacent to a transposon-associated motif.
16. The composition of claim 15, wherein the targeting sequence is 15-25 nucleotides in length.
17. The composition according to claim 15, wherein the transposon-associated motif is a nucleotide sequence represented by the following formula: TTTAA.
18. A recombinant vector, wherein the recombinant vector comprises a nucleic acid encoding the nuclease according to any one of claims 1-5, a reRNA according to claim 6, a guide RNA according to any one of claims 7-9, a nucleic acid according to claim 10, a nucleic acid construct according to any one of claims 11-14, or a composition according to any one of claims 15-17.
19. The recombinant vector according to claim 18, wherein the recombinant vector is selected from a recombinant cloning vector, a recombinant eukaryotic expression plasmid, or a recombinant viral vector.
20. The recombinant vector according to claim 19, wherein the recombinant eukaryotic expression plasmid is selected from pcDNA3.1, pCMV, pUC18, pUC19, pUC57, pBAD, pET, pENTR, pGenlenti, or pAAV.
21. The recombinant vector according to claim 19, wherein the recombinant viral vector is selected from a recombinant adenoviral vector, a recombinant adeno-associated viral vector, a recombinant retroviral vector, a recombinant herpes simplex viral vector, or a recombinant vaccinia viral vector.
22. A recombinant host cell, wherein the recombinant host cell comprises the nuclease according to any one of claims 1-5, the reRNA according to claim 6, the guide RNA according to any one of claims 7-9, the nucleic acid according to claim 10, the nucleic acid construct according to any one of claims 11-14, the composition according to any one of claims 15-17, or the recombinant vector according to any one of claims 18-21.
23. The recombinant host cell of claim 22, wherein the recombinant host cell is selected from an animal cell, a plant cell, a fungal cell, or a bacterial cell.
24. The recombinant host cell of claim 22, wherein the recombinant host cell is selected from an algae cell or a yeast cell.
25. The recombinant host cell of claim 23, wherein the animal cell is selected from a mammalian cell; wherein the plant cell is selected from a monocotyledonous plant cell or a dicotyledonous plant cell.
26. The recombinant host cell according to claim 25, wherein the mammalian cell is selected from primary cells, immortalized cell lines, embryonic stem cell lines and cells differentiated therefrom, or induced pluripotent stem cell lines and cells differentiated therefrom; wherein the monocotyledonous plant cell or dicotyledonous plant cell is selected from rice cells, corn cells or soybean cells.
27. The recombinant host cell of claim 25, wherein the mammalian cell is a cancer cell line.
28. The recombinant host cell of claim 26, wherein: The primary cells are selected from mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells; The immortalized cell line is selected from HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial cell line, hTERT-immortalized epithelial cell line, hTERT-immortalized fibroblast cell line, hTERT-immortalized keratinocyte cell line, and hTERT-immortalized ductal cell line; The embryonic stem cell line is selected from H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, and D3.
29. The recombinant host cell of claim 27, wherein: The cancer cell line is selected from Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, and Renca.
30. A method for introducing a double-strand break into a target gene in a host cell for non-diagnostic or therapeutic purposes, wherein the method comprises: The nuclease according to any one of claims 1 to 5, the reRNA according to claim 6, the guide RNA according to any one of claims 7 to 9, the nucleic acid according to claim 10, the nucleic acid construct according to any one of claims 11 to 14, the composition according to any one of claims 15 to 17, or the recombinant vector according to any one of claims 18 to 21 is delivered into a host cell.
31. A method for deleting, replacing or inserting a target gene in a host cell for non-diagnostic or therapeutic purposes, wherein the method comprises: The nuclease according to any one of claims 1 to 5, the reRNA according to claim 6, the guide RNA according to any one of claims 7 to 9, the nucleic acid according to claim 10, the nucleic acid construct according to any one of claims 11 to 14, the composition according to any one of claims 15 to 17, or the recombinant vector according to any one of claims 18 to 21 is delivered into a host cell.
32. A method for obtaining a host cell in which a target gene is deleted, replaced or inserted for non-diagnostic or therapeutic purposes, wherein the method comprises: The nuclease according to any one of claims 1 to 5, the reRNA according to claim 6, the guide RNA according to any one of claims 7 to 9, the nucleic acid according to claim 10, the nucleic acid construct according to any one of claims 11 to 14, the composition according to any one of claims 15 to 17, or the recombinant vector according to any one of claims 18 to 21 is delivered into a host cell.
33. The method of any one of claims 30-32, wherein the delivery method is selected from cationic liposome delivery, lipid nanoparticle delivery, cationic polymer delivery, vesicle-exosome delivery, gold nanoparticle delivery, polypeptide and protein delivery, retroviral delivery, lentiviral delivery, adenoviral delivery, adeno-associated virus delivery, electroporation, Agrobacterium infection, or gene gun.
34. The method of any one of claims 30-32, wherein the host cell is selected from an animal cell, a plant cell, a fungal cell, or a bacterial cell.
35. The method of any one of claims 30-32, wherein the host cell is selected from an algae cell or a yeast cell.
36. The method of claim 34, wherein the animal cell is selected from a mammalian cell; wherein the plant cell is selected from a monocotyledonous plant cell or a dicotyledonous plant cell.
37. The method according to claim 36, wherein the mammalian cell is selected from primary cells, immortalized cell lines, embryonic stem cell lines and cells differentiated therefrom, or induced pluripotent stem cell lines and cells differentiated therefrom; wherein the monocotyledonous plant cell or dicotyledonous plant cell is selected from rice cells, corn cells or soybean cells.
38. The method of claim 36, wherein the mammalian cell is a cancer cell line.
39. The method of claim 37, wherein: The primary cells are selected from mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells; The immortalized cell line is selected from HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial cell line, hTERT-immortalized epithelial cell line, hTERT-immortalized fibroblast cell line, hTERT-immortalized keratinocyte cell line, and hTERT-immortalized ductal cell line; The embryonic stem cell line is selected from H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, and D3.
40. The method of claim 38, wherein: The cancer cell line is selected from Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, Renca; 41. Use of the nuclease according to any one of claims 1-5, the reRNA according to claim 6, the guide RNA according to any one of claims 7-9, the nucleic acid according to claim 10, the nucleic acid construct according to any one of claims 11-14, the composition according to any one of claims 15-17, the recombinant vector according to any one of claims 18-21, or the recombinant host cell according to any one of claims 22-29 in the preparation of a drug or agent for introducing a double-strand break into a target gene of a host cell.
42. Use of the nuclease according to any one of claims 1-5, the reRNA according to claim 6, the guide RNA according to any one of claims 7-9, the nucleic acid according to claim 10, the nucleic acid construct according to any one of claims 11-14, the composition according to any one of claims 15-17, the recombinant vector according to any one of claims 18-21, or the recombinant host cell according to any one of claims 22-29 in the preparation of a drug or agent for deleting, replacing or inserting a target gene of a host cell.
43. The use according to any one of claims 41-42, wherein the host cell is selected from an animal cell, a plant cell, a fungal cell, or a bacterial cell.
44. The use according to any one of claims 41-42, wherein the host cell is selected from an algae cell or a yeast cell.
45. The use according to claim 43, wherein the animal cell is selected from mammalian cells; wherein the plant cell is selected from monocotyledonous plant cells or dicotyledonous plant cells.
46. The use according to claim 45, wherein the mammalian cell is selected from primary cells, immortalized cell lines, embryonic stem cell lines and cells differentiated therefrom, or induced pluripotent stem cell lines and cells differentiated therefrom; wherein the monocotyledonous plant cell or dicotyledonous plant cell is selected from rice cells, corn cells or soybean cells.
47. The use according to claim 45, wherein the mammalian cell is a cancer cell line.
48. The use according to claim 46, wherein: The primary cells are selected from mesenchymal stem cells, endothelial cells, epithelial cells, fibroblasts, keratinocytes, melanocytes, smooth muscle cells, and immune cells; The immortalized cell line is selected from HEK293, NIH-3T3, RAW-264.7, STO, VERO, CT26, hTERT-immortalized human endothelial cell line, hTERT-immortalized epithelial cell line, hTERT-immortalized fibroblast cell line, hTERT-immortalized keratinocyte cell line, and hTERT-immortalized ductal cell line; The embryonic stem cell line is selected from H1, H9, WIBR2, WIBR3, G-Olig2, ESF158, RW.4, R1, and D3.
49. The use according to claim 47, wherein: The cancer cell line is selected from Hela, HepG2 / 3, HL-60, HT-1080, HT-29, A549, SW620, HCT-15, HCT116, MDA-MB-231, MCF7, SK-OV-3, PANC-1, AsPc-1, THP-1, Huh7, KG-1, RAJI, HB-CB, Jurkat, K562, CRL5826, CHO, MDCK, and Renca.
50. Use of the nuclease according to any one of claims 1-5, the reRNA according to claim 6, the guide RNA according to any one of claims 7-9, the nucleic acid according to claim 10, the nucleic acid construct according to any one of claims 11-14, the composition according to any one of claims 15-17, the recombinant vector according to any one of claims 18-21, or the recombinant host cell according to any one of claims 22-29 in the preparation of a drug or preparation for gene therapy, cell therapy, genome research, stem cell induction and post-induction differentiation.
51. A kit, wherein the kit comprises a nuclease according to any one of claims 1-5, a reRNA according to claim 6, a guide RNA according to any one of claims 7-9, a nucleic acid according to claim 10, a nucleic acid construct according to any one of claims 11-14, a composition according to any one of claims 15-17, a recombinant vector according to any one of claims 18-21, or a recombinant host cell according to any one of claims 22-29.
Citation Information
Patent Citations
A novel RNA-programmable system for targeting polynucleotides
WO2023275601A1