Method for manufacturing DNA-edited eukaryotic cell and kit used in that method

The CRISPR-Cas3 system in eukaryotic cells uses pre-crRNA and a nuclear localization signal to overcome previous inefficiencies, enabling precise and extensive DNA editing by accurately targeting longer sequences and generating large deletions or insertions.

JP2025156592AActive Publication Date: 2025-10-14OSAKA UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025134235
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-06-08
Filing Date
2025-08-12
Publication Date
2025-10-14
Estimated Expiration
2038-06-08

AI Technical Summary

Technical Problem

Existing CRISPR-Cas3 systems have not been successfully established in eukaryotic cells for genome editing, despite their effectiveness in bacteria, due to challenges in using mature crRNA and inefficient DNA recognition and cleavage.

Method used

The introduction of a CRISPR-Cas3 system into eukaryotic cells using pre-crRNA and a nuclear localization signal, particularly a bipartite nuclear localization signal, enables efficient genome editing by recognizing target sequences more accurately and causing large deletions or insertions.

Benefits of technology

The CRISPR-Cas3 system achieves precise and extensive DNA editing in eukaryotic cells, including large deletions and insertions, surpassing the capabilities of CRISPR-Cas9 systems by recognizing longer sequences and targeting regions inaccessible to conventional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025156592000002
    Figure 2025156592000002
  • Figure 2025156592000003
    Figure 2025156592000003
  • Figure 2025156592000004
    Figure 2025156592000004
Patent Text Reader

Abstract

To establish a CRISPR-Cas3 system in a eukaryotic cell.SOLUTION: A method for manufacturing a DNA-edited eukaryotic cell comprises introducing a CRISPR-Cas3 system into a eukaryotic cell, where the CRISPR-Cas3 system includes the following (A)-(C): (A) a Cas3 protein, a polynucleotide encoding the protein, or an expression vector including the polynucleotide; (B) a cascade protein, a polynucleotide encoding the protein, or an expression vector including the polynucleotide; and (C) a crRNA, a polynucleotide encoding the crRNA, or an expression vector including the polynucleotide.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to methods for producing DNA-edited eukaryotic cells, animals, and plants, and kits for use in said methods. [Background technology]

[0002] Bacteria and archaea possess an adaptive immune system that specifically recognizes and eliminates foreign organisms, such as phages, that attempt to invade. This system, called the CRISPR-Cas system, first incorporates the genomic information of the foreign organism into its own genome (adaptation). Then, when the same foreign organism attempts to invade again, it uses the information incorporated into its own genome and the complementarity of the genomic sequence to cut and eliminate the foreign genome (interference).

[0003] Recently, genome editing (DNA editing) techniques have been developed using the above-mentioned CRISPR-Cas system as a "DNA editing tool" (Non-Patent Document 1).

[0004] CRISPR-Cas systems are broadly classified into "Class 1" systems, which consist of multiple Cass as effectors that work in the DNA cleavage process, and "Class 2" systems, which consist of a single Cas. Among Class 1 CRISPR-Cas systems, "Type I" systems involving Cas3 and the Cascade complex (meaning a complex of Cascade and crRNA; the same applies below) are widely known. Among Class 2 CRISPR-Cas systems, "Type II" systems involving Cas9 are widely known (hereinafter, with regard to CRISPR-Cas systems, "Class 1 Type I" and "Class 2 Type II" may be simply referred to as "Type I" and "Type II," respectively). Furthermore, Class 2 CRISPR-Cas systems involving Cas9 have been widely used in DNA editing technologies to date (hereinafter, sometimes referred to as the "CRISPR-Cas9 system"). For example, Non-Patent Document 1 reports a Class 2 CRISPR-Cas system that uses Cas9 to cleave DNA.

[0005] On the other hand, despite many efforts, no successful genome editing has been reported in eukaryotic cells using Class 1 CRISPR-Cas systems (hereinafter sometimes referred to as "CRISPR-Cas3 systems") that cleave DNA using Cas3 and a cascade complex. For example, Non-Patent Documents 2 and 3 simply report that the CRISPR-Cas3 system can completely degrade target DNA in a cell-free system and selectively eliminate specific E. coli strains. However, these reports do not imply successful genome editing, and no evidence has been provided in eukaryotic cells. Furthermore, Patent Document 1 proposes using FokI nuclease instead of Cas3 for genome editing in eukaryotic cells, since the CRISPR-Cas3 system degrades target DNA in E. coli due to the helicase and exonuclease activities of Cas3 (Example 5, Figure 6). Furthermore, Patent Document 2 proposes that the CRISPR-Cas3 system, which degrades target DNA in E. coli (Figure 4), be repurposed for programmable gene suppression by deleting Cas3 or using inactivated Cas3 (Cas3' and Cas3'') (e.g., Example 15, Claim 4(e)). [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Special Publication No. 2015-503535 [Patent Document 2] Special Publication No. 2017-512481 [Non-patent literature]

[0007] [Non-Patent Document 1] Jinek M et al. (2012) A Programmable Dual-RNA Guided DNA Endonuclease in Adaptive Bacterial Immunity, Science, Vol.337 (Issue 6096), pp.816-821 [Non-patent document 2] Mulepati S & Bailey S (2013) In Vitro Reconstitution of an Escherichia coli RNA-guided Immune System Reveals Unidirectional, ATP-dependent Degradation of DNA Target, Journal of Biological Chemistry, Vol.288 (No.31), pp.22184-22192 [Non-patent document 3] Ahmed A. Gomaa et al. (2014) Programmable Reomoval of Bacterial Strains by Use of Genome Targeting CRISPR-Cas Systems, mbio. asm. org, Volume 5, Issue 1, e00928-13 Summary of the Invention [Problem to be solved by the invention]

[0008] The present invention was made in light of these circumstances, and its purpose is to establish a CRISPR-Cas3 system in eukaryotic cells. [Means for solving the problem]

[0009] To achieve this goal, the inventors conducted extensive research and finally succeeded in establishing a CRISPR-Cas3 system in eukaryotic cells. The most widely used CRISPR-Cas9 system has been successfully used for genome editing in various eukaryotic cells, but this system typically uses mature crRNA as the crRNA. Surprisingly, however, genome editing in eukaryotic cells using mature crRNA with the CRISPR-Cas3 system was difficult. Effective genome editing was only possible with pre-crRNA, which is not typically used as a system component. Specifically, we found that cleavage of crRNA by proteins constituting the cascade is crucial for the CRISPR-Cas3 system to function in eukaryotic cells. This pre-crRNA-based CRISPR-Cas3 system was broadly applicable not only to type IE systems but also to type IF and type IG systems. Furthermore, the genome editing efficiency of the CRISPR-Cas3 system in eukaryotic cells was further improved by adding a nuclear localization signal, particularly the bipartite nuclear localization signal, to Cas3. The present inventors also discovered that the CRISPR-Cas3 system, unlike the CRISPR-Cas9 system, is capable of causing large deletions in regions including or upstream of the PAM sequence, and thus completed the present invention.

[0010] That is, the present invention relates to a CRISPR-Cas3 system in eukaryotic cells, and more specifically provides the following inventions.

[0011] [1] A method for producing a eukaryotic cell with edited DNA, comprising introducing a CRISPR-Cas3 system into the eukaryotic cell, wherein the CRISPR-Cas3 system comprises the following (A) to (C): (A) a Cas3 protein, a polynucleotide encoding the protein, or an expression vector containing the polynucleotide; (B) a Cascade protein, a polynucleotide encoding the protein, or an expression vector containing the polynucleotide; and (C) crRNA, a polynucleotide encoding the crRNA, or an expression vector containing the polynucleotide

[0012] [2] A method for producing an animal (excluding humans) or plant with edited DNA, comprising introducing a CRISPR-Cas3 system into the animal (excluding humans) or plant, wherein the CRISPR-Cas3 system comprises the following (A) to (C): (A) a Cas3 protein, a polynucleotide encoding the protein, or an expression vector containing the polynucleotide; (B) a Cascade protein, a polynucleotide encoding the protein, or an expression vector containing the polynucleotide; and (C) crRNA, a polynucleotide encoding the crRNA, or an expression vector containing the polynucleotide

[0013] [3] The method described in [1] or [2], which includes a step of introducing the CRISPR-Cas3 system into a eukaryotic cell and then cleaving the crRNA with a protein constituting the cascade protein.

[0014] [4] The method according to [1] or [2], wherein the crRNA is pre-crRNA.

[0015] [5] The method according to any one of [1] to [4], wherein a nuclear localization signal is added to the Cas3 protein and / or the Cascade protein.

[0016] [6] The method according to [5], wherein the nuclear localization signal is a bipartite nuclear localization signal.

[0017] [7] A kit for use in the method according to any one of [1] to [6], comprising the following (A) and (B): (A) a Cas3 protein, a polynucleotide encoding the protein, or an expression vector containing the polynucleotide; and (B) a cascade protein, a polynucleotide encoding the protein, or an expression vector containing the polynucleotide

[0018] [8] The kit described in [7], further comprising crRNA, a polynucleotide encoding the crRNA, or an expression vector containing the polynucleotide.

[0019] [9] The kit according to [8], wherein the crRNA is pre-crRNA.

[0020]

[10] The kit according to any one of [7] to [9], wherein a nuclear localization signal is added to the Cas3 protein and / or the cascade protein.

[0021]

[11] The kit according to

[10] , wherein the nuclear localization signal is a bipartite nuclear localization signal.

[0022] As used herein, the term "polynucleotide" refers to a polymer of nucleotides and is used synonymously with the terms "gene," "nucleic acid," or "nucleic acid molecule." Polynucleotides can exist in the form of DNA (e.g., cDNA or genomic DNA) or RNA (e.g., mRNA). The term "protein" is used synonymously with "peptide" or "polypeptide." [Effects of the Invention]

[0023] By using the CRISPR-Cas3 system of the present invention, it is now possible to edit DNA in eukaryotic cells. [Brief explanation of the drawings]

[0024] [Figure 1] This shows the results of an SSA assay that measured the cleavage activity against exogenous DNA. [Figure 2]FIG. 1 is a schematic diagram showing the location of the target sequence in the CCR5 gene. [Figure 3A] FIG. 1 shows a CCR5 gene (clone 1) in which part of the base sequence was deleted using the CRISPR-Cas3 system. [Figure 3B] FIG. 1 shows a CCR5 gene (clone 2) in which part of the base sequence was deleted using the CRISPR-Cas3 system. [Figure 3C] FIG. 1 shows a CCR5 gene (clone 3) in which part of the base sequence was deleted using the CRISPR-Cas3 system. [Figure 3D] FIG. 1 shows a CCR5 gene (clone 4) in which a portion of the base sequence was deleted using the CRISPR-Cas3 system. [Figure 4] (a) is a schematic diagram showing the structure of the Cascade plasmid. (b) is a schematic diagram showing the structure of the Cas3 plasmid. (c) is a schematic diagram showing the structure of the pre-crRNA plasmid. (d) is a schematic diagram showing the structure of the reporter vector (including the target sequence). [Figure 5] FIG. 1 is a schematic diagram showing the location of the target sequence in the EMX1 gene. [Figure 6A] This is a diagram showing the EMX1 gene (clone 1) in which part of the base sequence was deleted using the CRISPR-Cas3 system. [Figure 6B] This is a diagram showing the EMX1 gene (clone 2) in which another part of the base sequence was deleted using the CRISPR-Cas3 system. [Figure 7] FIG. 1 is a schematic diagram showing the structure of a Cas3 / Cascade plasmid with bpNLS added. [Figure 8] Schematic diagram showing the structure of the Cascade (2A) plasmid. [Figure 9] This shows the results of an SSA assay that measured the cleavage activity against exogenous DNA. [Figure 10A]Figure 1 shows the structures of the pre-crRNA (LRSR and RSR) and mature crRNA used in this example. The underlined part in the figure indicates the 5' handle (Cas5 handle), and the double underlined part indicates the 3' handle (Cas6 handle). [Figure 10B] FIG. 1 shows the results of SSA assay using pre-crRNA (LRSR and RSR) and mature crRNA. [Figure 11] FIG. 1 shows the results of SSA assays using one NLS or two NLSs (bpNLS) in plasmids for expression of Cas3 / Cascade genes. [Figure 12] FIG. 1 shows the effect of PAM sequences on the DNA cleavage activity of the CRISPR-Cas3 system. [Figure 13] FIG. 1 shows the effect of a single mismatch in the spacer on the DNA cleavage activity of the CRISPR-Cas3 system. [Figure 14] FIG. 1 shows the effect of Cas3 mutations in the HD nuclease domain (H74A), SF2 helicase domain motif 1 (K320A), and motif 3 (S483 / T485A). [Figure 15] FIG. 1 shows a comparison of the DNA cleavage activity of type IE, type IF, and type IG CRISPR-Cas3 systems. [Figure 16] This figure shows the size of deletions caused by the CRISPR-Cas3 system, detected by sequencing TA cloning samples of PCR products. [Figure 17] Figure 1 shows the locations of deletions caused by the CRISPR-Cas3 system detected by high-throughput sequencing of TA clones (n=49). [Figure 18A] This figure shows the number of deletions detected by the CRISPR-Cas3 system per deletion size using microarray-based capture sequencing of over 1000 kb surrounding the targeted EMX1 locus. [Figure 18B]This figure shows the number of deletions detected by the CRISPR-Cas3 system per deletion size using microarray-based capture sequencing of over 1000 kb surrounding the targeted CCR5 locus. DETAILED DESCRIPTION OF THE INVENTION

[0025] [1] Methods for producing DNA-edited eukaryotic cells, animals, and plants The method of the present invention comprises introducing a CRISPR-Cas3 system into a eukaryotic cell, wherein the CRISPR-Cas3 system comprises the following (A) to (C): (A) a Cas3 protein, a polynucleotide encoding the protein, or an expression vector containing the polynucleotide; (B) a Cascade protein, a polynucleotide encoding the protein, or an expression vector containing the polynucleotide; and (C) crRNA, a polynucleotide encoding the crRNA, or an expression vector containing the polynucleotide

[0026] Class 1 CRISPR-Cas systems are classified into Type I and Type III. Type I systems are further subdivided into six types, Type IA, Type IB, Type IC, Type ID, Type IE, and Type IF, as well as Type IG, a subtype of Type IB, depending on the type of proteins that make up the cascade (hereinafter simply referred to as "cascade" or "cascade proteins"). (See, for example, [van der Oost J et al. (2014) Unravelling the structural and mechanistic basis of CRISPR-Cas systems, Nature Reviews Microbiologym, Vol. 12 (No. 7), pp. 479-492] and [Jackson RN et al. (2014) Fitting CRISPR-associated Cas3 into the Helicase Family Tree, Current Opinion in Structural Biology, Vol. 24, pp. 106-114].

[0027] The Type I CRISPR-Cas system functions by cleaving DNA through the cooperation of Cas3 (a protein with nuclease and helicase activity), Cascade, and crRNA. Because Cas3 is used as a nuclease, this system is referred to as the "CRISPR-Cas3 system" in this invention.

[0028] Use of the CRISPR-Cas3 system of the present invention provides the following advantages, for example:

[0029] First, the crRNA used in the CRISPR-Cas3 system generally recognizes target sequences of 32–37 bases (Ming Li et al., Nucleic Acids Res. 2017 May 5; 45(8): 4642–4654). In contrast, the crRNA used in the CRISPR-Cas9 system generally recognizes target sequences of 18–24 bases. Therefore, the CRISPR-Cas3 system is thought to be able to recognize target sequences more accurately than the CRISPR-Cas9 system.

[0030] Furthermore, the PAM sequence of the CRISPR-Cas9 system, a Class 2 Type II system, is "NGG (N is any base)" adjacent to the 3' side of the target sequence. Furthermore, the PAM sequence of the CRISPR-Cpf1 system, a Class 2 Type V system, is "AA" adjacent to the 5' side of the target sequence. In contrast, the PAM sequence of the CRISPR-Cas3 system of the present invention is "AAG" or a base sequence similar thereto (e.g., "AGG," "GAG," "TAC," "ATG," "TAG," etc.) adjacent to the 5' side of the target sequence (Figure 12). Therefore, it is believed that the CRISPR-Cas3 system of the present invention can be used to target DNA editing in regions that could not be recognized by conventional methods.

[0031] Furthermore, unlike the Class 2 CRISPR-Cas systems, the CRISPR-Cas3 system generates DNA breaks at multiple sites. Therefore, the CRISPR-Cas3 system of the present invention can generate extensive deletion mutations of hundreds to thousands of bases, or even more in some cases (Figures 3, 6, 16-18). This function is believed to enable the system to be used to knock out long genomic regions or knock in long DNA. When performing knock-in, donor DNA is typically used, and this donor DNA also serves as a molecule constituting the CRISPR-Cas3 system of the present invention.

[0032] In this specification, when simply referred to as "Cas3," it means "Cas3 protein." The same applies to Cascade proteins.

[0033] The CRISPR-Cas3 system of the present invention encompasses all six subtypes of Type I. That is, although the proteins that make up the CRISPR-Cas3 system may differ slightly in composition depending on the subtype (e.g., the proteins that make up the cascade differ), the present invention encompasses all of these proteins. Indeed, in this example, it was found that genome editing is possible not only in Type IE, but also in Type 1-G and Type IF systems (Figure 15).

[0034] Type IE CRISPR-Cas3 systems, which are the most common among type I CRISPR-Cas3 systems, cleave DNA by the cooperation of crRNA with Cas3 and Cascade (Cse1 (Cas8), Cse2 (Cas11), Cas5, Cas6, and Cas7).

[0035] In the type IA system, the cascade consists of Cas8a1, Csa5 (Cas11), Cas5, Cas6, and Cas7; in the type IB system, the cascade consists of Cas8b1, Cas5, Cas6, and Cas7; in the type IC system, the cascade consists of Cas8c, Cas5, and Cas7; in the type ID system, the cascade consists of Cas10d, Csc1 (Cas5), Cas6, and Csc2 (Cas7); in the type IF system, the cascade consists of Csy1 (Cas8f), Csy2 (Cas5), Cas6, and Csy3 (Cas7); and in the type IG system, the cascade consists of Cst1 (Cas8a1), Cas5, Cas6, and Cst2 (Cas7). In the present invention, Cas3 and Cascade are collectively referred to as the "Cas protein group."

[0036] Below, we will explain the type IE CRISPR-Cas3 system as a representative example, but for other types of CRISPR-Cas3 systems, the cascades that make up the system can be interpreted as appropriate.

[0037] -Cas protein group- In the CRISPR-Cas3 system of the present invention, the Cas proteins can be introduced into eukaryotic cells in the form of proteins, polynucleotides encoding the proteins, or expression vectors containing the polynucleotides. Introducing the Cas proteins into eukaryotic cells in the form of proteins allows for the appropriate adjustment of the amount of each protein, which is advantageous from the standpoint of ease of handling. Furthermore, taking into account factors such as intracellular cleavage efficiency, a complex of the Cas proteins can be formed first and then introduced into eukaryotic cells.

[0038] In the present invention, it is preferable to add a nuclear localization signal to the Cas proteins. The nuclear localization signal can be added to the N-terminus and / or C-terminus of the Cas proteins (the 5'-terminus and / or 3'-terminus of the polynucleotide encoding each Cas protein). Adding a nuclear localization signal to the Cas proteins in this way promotes their localization to the nucleus within the cell, resulting in the advantage of efficient DNA editing.

[0039] The nuclear localization signal is a peptide sequence consisting of several to several tens of basic amino acids, and the sequence is not particularly limited as long as it allows a protein to be transported into the nucleus. Specific examples of such nuclear localization signals are described in, for example, Wu J et al. (2009) The Intracellular Mobility of Nuclear Import Receptors and NLS Cargoes, Biophysical Journal, Vol. 96 (Issue 9), pp. 3840-3849. Any nuclear localization signal commonly used in the art can be used in the present invention.

[0040] The nuclear localization signal may be, for example, PKKKRKV (SEQ ID NO: 52) (encoded by the nucleotide sequence CCCAAGAAGAAGCGGAAGGTG (SEQ ID NO: 53)). When using the nuclear localization signal, it is preferable to place a polynucleotide consisting of the nucleotide sequence of SEQ ID NO: 53 at the 5' end of each polynucleotide encoding the Cas proteins. Alternatively, the nuclear localization signal may be, for example, KRTADGSEFESPKKKRKVE (SEQ ID NO: 54) (encoded by the nucleotide sequence AAGCGGACTGCTGATGGCAGTGAATTTGAGTCCCCAAAGAAGAAGAGAAAGGTGGAA (SEQ ID NO: 55)). When using the nuclear localization signal, it is preferable to place a polynucleotide consisting of the nucleotide sequence of SEQ ID NO: 55 at both ends of each polynucleotide encoding the Cas proteins (i.e., use of a "bipartite nuclear localization signal (bpNLS)").

[0041] These modifications, together with the use of pre-crRNA described below, are important for the efficient expression and function of the CRISPR-Cas3 system of the present invention in eukaryotic cells.

[0042] One preferred embodiment of the Cas protein group used in the present invention is as follows. Cas3: a protein encoded by a polynucleotide consisting of the nucleotide sequence shown in SEQ ID NO: 1 or SEQ ID NO: 7 Cse1 (Cas8): A protein encoded by a polynucleotide consisting of the nucleotide sequence shown in SEQ ID NO: 2 or SEQ ID NO: 8 Cse2 (Ca11); a protein encoded by a polynucleotide consisting of the nucleotide sequence shown in SEQ ID NO: 3 or SEQ ID NO: 9 Cas5: a protein encoded by a polynucleotide consisting of the nucleotide sequence shown in SEQ ID NO: 4 or SEQ ID NO: 10 Cas6: a protein encoded by a polynucleotide consisting of the nucleotide sequence shown in SEQ ID NO: 5 or SEQ ID NO: 11 Cas7: a protein encoded by a polynucleotide consisting of the nucleotide sequence shown in SEQ ID NO: 6 or SEQ ID NO: 12 The Cas proteins are (1) proteins in which the nuclear localization signal PKKKRKV (SEQ ID NO: 52) has been added to the N-terminus of wild-type E. coli Cas3, Cse1 (Cas8), Cse2 (Cas11), Cas5, Cas6, and Cas7, or (2) proteins in which the nuclear localization signal KRTADGSEFESPKKKRKVE (SEQ ID NO: 54) has been added to the N- and C-termini of wild-type E. coli Cas3, Cse1 (Cas8), Cse2 (Cas11), Cas5, Cas6, and Cas7. Proteins with these amino acid sequences can be translocated into the nucleus of eukaryotic cells. Once translocated into the nucleus, the Cas proteins cleave the target DNA. Furthermore, editing of target DNA is possible even in regions of DNA with rigid structures (e.g., heterochromatin), which is thought to be difficult with the CRISPAR-Cas9 system.

[0043] Another embodiment of each protein of the Cas protein group used in the present invention is a protein encoded by a nucleotide sequence having 90% or more sequence identity with the nucleotide sequence of the above-mentioned Cas protein group. Another embodiment of each protein of the Cas protein group used in the present invention is a protein encoded by a polynucleotide that hybridizes under stringent conditions to a polynucleotide consisting of a nucleotide sequence complementary to the nucleotide sequence of the above-mentioned Cas protein group. Each of the above proteins has DNA cleavage activity when it forms a complex with another protein that constitutes the Cas protein group. The meanings of terms such as "sequence identity" and "stringent conditions" will be explained below.

[0044] -Polynucleotide encoding the Cas protein group- Polynucleotides encoding wild-type proteins constituting the type IE CRISPR-Cas system include polynucleotides modified for efficient expression in eukaryotic cells. That is, modified polynucleotides encoding the Cas proteins can be used. One preferred embodiment of polynucleotide modification is modification to a nucleotide sequence suitable for expression in eukaryotic cells, for example, by optimizing codons for expression in eukaryotic cells.

[0045] One preferred embodiment of the polynucleotide encoding the Cas protein group used in the present invention is as follows. Cas3: a polynucleotide consisting of the nucleotide sequence shown in SEQ ID NO: 1 or SEQ ID NO: 7 Cse1 (Cas8): a polynucleotide consisting of the nucleotide sequence shown in SEQ ID NO: 2 or SEQ ID NO: 8 Cse2 (Ca11): a polynucleotide consisting of the nucleotide sequence shown in SEQ ID NO: 3 or SEQ ID NO: 9 Cas5: a polynucleotide consisting of the nucleotide sequence shown in SEQ ID NO: 4 or SEQ ID NO: 10 Cas6: a polynucleotide consisting of the nucleotide sequence shown in SEQ ID NO: 5 or SEQ ID NO: 11 Cas7: a polynucleotide consisting of the nucleotide sequence shown in SEQ ID NO: 6 or SEQ ID NO: 12 These are polynucleotides that have been artificially modified from the base sequences encoding the wild-type Cas proteins of E. coli (Cas3; SEQ ID NO: 13, Cse1 (Cas8); SEQ ID NO: 14, Cse2 (Cas11); SEQ ID NO: 15, Cas5; SEQ ID NO: 16, Cas6; SEQ ID NO: 17, Cas7; SEQ ID NO: 18) so that they can be expressed and function in mammalian cells.

[0046] The artificial modification of the polynucleotide involves modifying the base sequence to one suitable for expression in eukaryotic cells and adding a nuclear localization signal. The modification of the base sequence and the addition of the nuclear localization signal are as described above. This is expected to result in a more sufficient increase in the expression level and function of the Cas proteins.

[0047] Another embodiment of the polynucleotides encoding the Cas proteins used in the present invention is a polynucleotide obtained by modifying a nucleotide sequence encoding a wild-type Cas protein group, the nucleotide sequence of which has 90% or more sequence identity to the nucleotide sequence of the above-mentioned Cas protein group. The proteins expressed from these polynucleotides have DNA cleavage activity when they form a complex with proteins expressed from other polynucleotides that constitute the Cas protein group.

[0048] The sequence identity of the nucleotide sequences may be at least 90% or more, more preferably 95% or more (e.g., 95%, 96%, 97%, 98%, 99% or more), over the entire nucleotide sequence (or the region encoding the portion necessary for Cse3 function). The identity of the nucleotide sequences can be determined using a program such as BLASTN (see Altschul SF (1990) Basic local alignment search tool, Journal of Molecular Biology, Vol. 215 (Issue 3), pp. 403-410). An example of parameters for analyzing nucleotide sequences using BLASTN is a score of 100 and a word length of 12. Specific techniques for performing BLASTN analysis are known to those skilled in the art. Additions or deletions (e.g., gaps) may be allowed to optimally align the nucleotide sequences to be compared.

[0049] Furthermore, the phrase "having DNA cleavage activity" means being able to cleave a polynucleotide chain at at least one site.

[0050] The CRISPR-Cas3 system of the present invention preferably specifically recognizes a target sequence and cleaves DNA. Whether the CRISPR-Cas3 system specifically recognizes a target sequence can be determined, for example, by the dual-luciferase assay described in Example A-1.

[0051] Another embodiment of the polynucleotides encoding the Cas proteins used in the present invention is a polynucleotide that hybridizes under stringent conditions with a polynucleotide consisting of a nucleotide sequence complementary to the nucleotide sequence of the above-mentioned Cas proteins. The proteins expressed from these polynucleotides have DNA cleavage activity when they form a complex with proteins expressed from other polynucleotides that constitute the Cas proteins.

[0052] Here, "stringent conditions" refers to conditions under which two polynucleotide chains form a double-stranded polynucleotide specific to their base sequences, but do not form a non-specific double-stranded polynucleotide. In other words, "hybridizing under stringent conditions" can be said to be conditions under which hybridization can occur within a temperature range of 15°C lower, preferably 10°C lower, and more preferably 5°C lower than the melting temperature (Tm value) of nucleic acids with high sequence identity (e.g., a perfectly matched hybrid).

[0053] An example of stringent conditions is as follows: First, two types of polynucleotides are hybridized in a buffer solution (pH 7.2) consisting of 0.25 M NaHPO, 7% SDS, 1 mM EDTA, and 1x Denhardt's solution at 60 to 68°C (preferably 65°C, more preferably 68°C) for 16 to 24 hours. Then, the hybridization is performed twice for 15 minutes at 60 to 68°C (preferably 65°C, more preferably 68°C) in a buffer solution (pH 7.2) consisting of 20 mM NaHPO, 1% SDS, and 1 mM EDTA.

[0054] Another example is the following method: First, prehybridization is performed overnight at 42°C in a hybridization solution containing 25% formamide (50% formamide for more stringent conditions), 4x SSC (sodium chloride / sodium citrate), 50 mM Hepes (pH 7.0), 10x Denhardt's solution, and 20 μg / mL denatured salmon sperm DNA. After that, a labeled probe is added and the mixture is incubated overnight at 42°C to hybridize the two polynucleotides.

[0055] Next, wash under one of the following conditions: Normal conditions: Wash at approximately 37°C using 1x SSC and 0.1% SDS as the wash solution; Stringent conditions: Wash at approximately 42°C using 5x SSC and 0.1% SDS as the wash solution; Even more stringent conditions: Wash at approximately 65°C using 0.2x SSC and 0.1% SDS as the wash solution.

[0056] Thus, the stricter the washing conditions for hybridization, the more specific the hybridization. The above-mentioned combinations of SSC, SDS, and temperature conditions are merely examples. Similar stringencies can be achieved by appropriately combining the above factors that determine hybridization stringency or other factors (e.g., probe concentration, probe length, hybridization reaction time, etc.). This is described, for example, in [Joseph Sambrook & David W. Russell, Molecular cloning: a laboratory manual, 3rd Ed., New York: Cold Spring Harbor Laboratory Press, 2001].

[0057] -Expression vector containing a polynucleotide encoding a Cas protein group- In the present invention, an expression vector can be used to express a Cas protein group. Various commonly used vectors can be used as base vectors for the expression vector, and the vector can be appropriately selected depending on the cell to be introduced or the method of introduction. Specifically, a plasmid, phage, cosmid, etc. can be used. The specific type of vector is not particularly limited, and a vector that can be expressed in a host cell can be appropriately selected.

[0058] Examples of the above-mentioned expression vectors include phage vectors, plasmid vectors, viral vectors, retroviral vectors, chromosomal vectors, episomal vectors and virus-derived vectors (such as bacterial plasmids, bacteriophages, yeast episomes), yeast chromosomal elements and viruses (such as baculoviruses, papovaviruses, vaccinia viruses, adenoviruses, avian poxviruses, pseudorabies viruses, herpes viruses, lentiviruses, retroviruses), and vectors derived from combinations thereof (such as cosmids and phagemids).

[0059] Expression vectors further contain sites for transcription initiation and termination, and preferably contain a ribosome binding site in the transcribed region. The coding portion of the mature transcript in the vector will contain the translation initiation codon AUG at the beginning of the polypeptide to be translated and a termination codon appropriately positioned at the end.

[0060] In the present invention, the expression vector for expressing the Cas proteins may contain a promoter sequence. The promoter sequence may be appropriately selected depending on the type of eukaryotic cell used as the host. The expression vector may also contain a sequence for enhancing transcription from DNA, such as an enhancer sequence. Examples of enhancers include the SV40 enhancer (located 100 to 270 bp downstream of the replication origin), the cytomegalovirus early promoter enhancer, and the polyoma and adenovirus enhancers located downstream of the replication origin. The expression vector may also contain a sequence for stabilizing the transcribed RNA, such as a polyA addition sequence (polyadenylation sequence, polyA). Examples of polyA addition sequences include the polyA addition sequence derived from the growth hormone gene, the polyA addition sequence derived from the bovine growth hormone gene, the polyA addition sequence derived from the human growth hormone gene, the polyA addition sequence derived from the SV40 virus, and the polyA addition sequence derived from the human or rabbit β-globin gene.

[0061] The number of polynucleotides encoding Cas proteins incorporated into the same vector is not particularly limited, as long as the CRISPR-Cas system can function in a host cell into which the expression vector has been introduced. For example, it is possible to design the polynucleotides encoding the Cas proteins to be incorporated into a single (same) vector, or it is also possible to design the polynucleotides encoding each Cas protein to be incorporated into separate vectors. For example, it is possible to design the polynucleotides encoding the Cascade proteins to be incorporated into a single (same) vector, and the polynucleotide encoding Cas3 to be incorporated into another vector. Preferably, from the standpoint of expression efficiency, etc., a method is used in which the polynucleotides encoding each Cas protein are incorporated into six different vectors.

[0062] Additionally, multiple polynucleotides encoding the same protein may be carried in the same vector for purposes such as adjusting the expression level, etc. For example, a design is possible in which a polynucleotide encoding Cas3 is placed at two locations within one type (same) vector.

[0063] Alternatively, an expression vector may be used that contains multiple nucleotide sequences encoding Cas proteins, with a nucleotide sequence encoding an amino acid sequence (e.g., 2A peptide) cleaved by intracellular proteases inserted between the multiple nucleotide sequences (see, for example, the vector structure in Figure 8). When a polynucleotide containing such a nucleotide sequence is transcribed and translated, a single linked polypeptide chain is expressed within the cell. Subsequently, the Cas proteins are separated by the action of intracellular proteases, forming individual proteins that then function as a complex. This allows for the quantitative ratio of Cas proteins expressed within the cell to be adjusted. For example, an expression vector containing one nucleotide sequence encoding Cas3 and one encoding Cse1 (Cas8) is predicted to express equal amounts of Cas3 and Cse1 (Cas8). Furthermore, the ability to express multiple Cas proteins using a single expression vector offers the advantage of ease of handling. However, from the perspective of high DNA cleavage activity, it is generally preferable to express each Cas protein using a different expression vector.

[0064] The expression vector used in the present invention can be prepared by known techniques. These techniques include those described in the operating manuals that accompany kits for preparing vectors, as well as those described in various manuals. For example, Joseph Sambrook & David W. Russell, Molecular cloning: a laboratory manual, 3rd Ed., New York: Cold Spring Harbor Laboratory Press, 2001, is a comprehensive manual.

[0065] -crRNA, a polynucleotide encoding the crRNA, or an expression vector containing the polynucleotide- The CRISPR-Cas3 system of the present invention comprises a crRNA, a polynucleotide encoding the crRNA, or an expression vector containing the polynucleotide for targeting DNA to perform genome editing.

[0066] crRNA is an RNA that forms part of the CRISPR-Cas system and has a base sequence complementary to the target sequence. The CRISPR-Cas3 system of the present invention uses crRNA to specifically recognize and cleave the target sequence. In CRISPR-Cas systems, such as the CRISPR-Cas9 system, mature crRNA has typically been used as the cRNA. However, it has become clear that the use of mature crRNA is not suitable for the functioning of the CRISPR-Cas3 system in eukaryotic cells, although the reason for this is unclear. Surprisingly, it has been found that the use of pre-crRNA instead of mature crRNA enables highly efficient genome editing in eukaryotic cells. This fact is evident from a comparison experiment between mature crRNA and pre-crRNA (Figure 10). Therefore, it is particularly preferable to use pre-crRNA as the crRNA in the present invention.

[0067] The pre-crRNA used in the present invention typically has a "leader sequence-repeat sequence-spacer sequence-repeat sequence (LRSR structure)" or "repeat sequence-spacer sequence-repeat sequence (RSR structure)" structure. The leader sequence is an AT-rich sequence and functions as a promoter for expressing the pre-crRNA. The repeat sequence is a sequence repeated via a spacer sequence, which is designed in the present invention as a sequence complementary to the target DNA (originally a sequence derived from foreign DNA incorporated during the adaptation process). The pre-crRNA is cleaved into mature crRNA by proteins constituting the cascade (e.g., Cas6 in types IA, B, D-E, and Cas5 in type IC).

[0068] Typically, the leader sequence is 86 bases long, and the repeat sequence is 29 bases long. The spacer sequence is, for example, 10 to 60 bases long, preferably 20 to 50 bases long, more preferably 25 to 40 bases long, and typically 32 to 37 bases long. Therefore, the length of the pre-crRNA used in the present invention is, for example, 154 to 204 bases long, preferably 164 to 194 bases long, more preferably 169 to 184 bases long, and typically 176 to 181 bases long in the case of an LRSR structure. Furthermore, the length is, for example, 68 to 118 bases long, preferably 78 to 108 bases long, more preferably 83 to 98 bases long, and typically 90 to 95 bases long in the case of an RSR structure.

[0069] For the CRISPR-Cas3 system of the present invention to function in eukaryotic cells, the process by which the repeat sequence of the pre-crRNA is cleaved by proteins constituting the cascade is believed to be important. Therefore, it should be understood that the repeat sequence may be shorter or longer than the above-mentioned chain length, as long as such cleavage occurs. In other words, the pre-crRNA can be said to be a crRNA in which sequences sufficient for cleavage by proteins constituting the cascade are added to both ends of the mature crRNA described below. A preferred embodiment of the method of the present invention thus includes a step in which the crRNA is cleaved by proteins constituting the cascade after introducing the CRISPR-Cas3 system into eukaryotic cells.

[0070] On the other hand, the mature crRNA generated by cleavage of pre-crRNA has a structure of "5' handle sequence-spacer sequence-3' handle sequence." Typically, the 5' handle sequence consists of 8 bases, from positions 22 to 29, of the repeat sequence and is held by Cas5. Typically, the 3' handle sequence consists of 21 bases, from positions 1 to 21, of the repeat sequence, forming a stem-loop structure between positions 6 and 21, which is held by Cas6. Therefore, the length of the mature crRNA is usually 61 to 66 bases. However, depending on the type of CRISPR-Cas3 system, there are mature crRNAs that do not have a 3' handle sequence, and in this case, the length is shortened by 21 bases.

[0071] The sequence of the RNA may be appropriately designed depending on the target sequence for which DNA editing is desired. RNA synthesis may be performed using any method known in the art.

[0072] -Eukaryotic cells- "Eukaryotic cells" in the present invention include, for example, animal cells, plant cells, algae cells, and fungal cells. Animal cells include, for example, mammalian cells as well as cells of fish, birds, reptiles, amphibians, and insects.

[0073] "Animal cells" include, for example, cells constituting an individual animal, cells constituting organs or tissues extracted from an animal, and cultured cells derived from animal tissues. Specific examples include germ cells such as oocytes and sperm; germ cells of various stages of embryos (e.g., 1-cell, 2-cell, 4-cell, 8-cell, 16-cell, and morula stages); stem cells such as induced pluripotent stem (iPS) cells and embryonic stem (ES) cells; and somatic cells such as fibroblasts, hematopoietic cells, neurons, muscle cells, bone cells, liver cells, pancreatic cells, brain cells, and kidney cells. Pre- and post-fertilization oocytes can be used as oocytes for creating genome-edited animals, but fertilized oocytes, i.e., fertilized eggs, are preferred. Pronuclear stage fertilized eggs are particularly preferred. Oocytes can be used by thawing cryopreserved oocytes.

[0074] In the present invention, the term "mammal" encompasses both humans and non-human mammals. Examples of non-human mammals include ungulates such as cattle, boars, pigs, sheep, and goats, perissodactyls such as horses, rodents such as mice, rats, guinea pigs, hamsters, and squirrels, lagomorphs such as rabbits, and carnivores such as dogs, cats, and ferrets. The non-human mammals may be livestock or companion animals (pets), or wild animals.

[0075] "Plant cells" include, for example, cells of grains, oil crops, forage crops, fruits, and vegetables. "Plant cells" include, for example, cells that constitute an individual plant, cells that constitute organs or tissues separated from a plant, and cultured cells derived from plant tissue. Examples of plant organs and tissues include leaves, stems, shoot tips (growing points), roots, tubers, and callus. Examples of plants include rice, corn, banana, peanut, sunflower, tomato, rapeseed, tobacco, wheat, barley, potato, soybean, cotton, and carnation, as well as their propagation materials (e.g., seeds, tuberous roots, tubers, etc.).

[0076] -DNA editing- In the present invention, "editing the DNA of a eukaryotic cell" may refer to a process of editing the DNA of a eukaryotic cell in vivo or in vitro. Furthermore, "editing DNA" refers to the following types of manipulations (including combinations thereof):

[0077] In this specification, DNA as used in the above context includes not only DNA present in the cell nucleus, but also DNA present outside the cell nucleus, such as mitochondrial DNA, and exogenous DNA. 1. Cleavage of the DNA strand at the target site. 2. Delete bases in the DNA strand at the target site. 3. Inserting bases into the DNA strand at the target site. 4. Substitute bases in the DNA strand at the target site. 5. Modify the bases of the DNA strand at the target site. 6. Regulates transcription of DNA (genes) at target sites.

[0078] One embodiment of the CRISPR-Cas3 system of the present invention utilizes a protein with enzymatic activity that modifies target DNA by a method other than introducing DNA breaks. This embodiment can be achieved, for example, by fusing Cas3 or Cascade with a heterologous protein having the desired enzymatic activity to form a chimeric protein. Therefore, the terms "Cas3" and "Cascade" of the present invention also include such fusion proteins. Examples of enzymatic activities of the fused protein include, but are not limited to, deaminase activity (e.g., cytidine deaminase activity, adenosine deaminase activity), methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, photolyase activity, and glycosylase activity. In this case, the nuclease activity and helicase activity of Cas3 are not necessarily required, and therefore, mutants lacking some or all of these activities (e.g., D domain H74A mutant (dnCas3), SF2 domain motif 1 K320N mutant (dhCas3), and SF2 domain motif 3 S483A / T485A double mutant (dh2Cas3)) can be used as Cas3. For example, by using a fusion protein of a deaminase with a mutant lacking some or all of the nuclease activity of Cas3 as a component of the CRISPR-Cas3 system of the present invention, precise genome editing can be achieved by replacing bases without causing large deletions at the target site. Methods for applying deaminases to CRISPR-Cas systems are known (Nishida K. et al., Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems, Science, DOI: 10.1126 / science.aaf8729, (2016)), and can be applied to the CRISPR-Cas3 system of the present invention.

[0079] In another embodiment of the CRISPR-Cas3 system of the present invention, transcription of a gene at the binding site of the system is regulated without DNA cleavage. This can be achieved, for example, by fusing Cas3 or Cascade with a desired transcriptional regulatory protein to form a chimeric protein. Therefore, the terms "Cas3" and "Cascade" as used herein also encompass such fusion proteins. Examples of transcriptional regulatory proteins include, but are not limited to, light-inducible transcriptional regulators, small molecule / drug-responsive transcriptional regulators, transcription factors, and transcriptional repressors. In this case, the nuclease and helicase activities of Cas3 are not necessarily required. Therefore, Cas3 mutants lacking some or all of these activities can be used (e.g., a D domain H74A mutant (dnCas3), a SF2 domain motif 1 K320N mutant (dhCas3), and a SF2 domain motif 3 S483A / T485A double mutant (dh2Cas3)). Techniques for applying transcriptional regulatory proteins to CRISPR-Cas systems are known to those skilled in the art.

[0080] Furthermore, in the CRISPR-Cas3 system of the present invention, for example, when a mutant in which the nuclease activity of Cas3 is partially or completely deleted is used, another protein having nuclease activity may be fused to Cas3 or the Cascade. Such embodiments are also included in the present invention.

[0081] In addition, in the CRISPR-Cas3 system of the present invention, when a mutant in which part or all of the nuclease activity of Cas3 is deleted is used and the activity of another protein is used in DNA editing, the term "DNA cleavage activity" in this specification shall be interpreted as various activities possessed by the other protein, as appropriate.

[0082] Furthermore, DNA editing may be performed on DNA contained in specific cells within an individual, for example, by targeting specific cells among the cells that make up an individual animal or plant.

[0083] The method for introducing the molecules constituting the CRISPR-Cas3 system of the present invention into eukaryotic cells in the form of a polynucleotide or an expression vector containing the polynucleotide is not particularly limited. Examples include electroporation, calcium phosphate transfection, liposome transfection, DEAE-dextran transfection, microinjection, cationic lipid-mediated transfection, electroporation, transduction, and infection with a viral vector. Such methods are described in many standard laboratory manuals, such as "Leonard G. Davis et al., Basic Methods in Molecular Biology, New York: Elsevier, 1986."

[0084] The method for introducing the CRISPR-Cas3 system of the present invention into eukaryotic cells in the form of a molecule protein is not particularly limited, and examples thereof include electroporation, cationic lipid-mediated transfection, and microinjection.

[0085] DNA editing according to the present invention can be applied to a variety of fields, including, for example, gene therapy, breeding, production of transgenic animals or cells, production of useful substances, and life science research.

[0086] Methods for producing non-human individuals from cells can be based on known methods. When producing non-human individuals from cells in animals, germ cells or pluripotent stem cells are typically used. For example, molecules constituting the CRISPR-Cas3 system of the present invention are introduced into oocytes, and the resulting oocytes are then implanted into the uterus of a pseudopregnant female non-human mammal, followed by the production of offspring. Implantation can be performed using fertilized eggs at the 1-cell, 2-cell, 4-cell, 8-cell, 16-cell, or morula stage. Oocytes can be cultured under appropriate conditions until implantation, if necessary. Oocyte implantation and culture can be performed using conventionally known techniques (Nagy A. et al., Manipulating the Mouse Embryo. Cold Spring Harbour, New York: Cold Spring Harbour Laboratory Press, 2003). From the resulting non-human individuals, offspring or clones with desired DNA editing can also be obtained.

[0087] Furthermore, it has long been known that somatic cells of plants possess totipotency, and methods for regenerating plants from plant cells have been established for various plants. Therefore, for example, by introducing the molecules constituting the CRISPR-Cas3 system of the present invention into plant cells and regenerating plants from the resulting plant cells, it is possible to obtain plants with the desired DNA knocked in. From the resulting plants, it is also possible to obtain offspring, clones, or propagation materials with the desired DNA edited. Methods established in the art can be used to regenerate plant tissues by tissue culture to obtain individuals (Transformation Protocols [Plant Edition], edited by Yutaka Tabei, Kagaku Dojin, pp. 340-347 (2012)).

[0088] [2] Kits used in the CRISPR-Cas3 system A kit for use in the CRISPR-Cas3 system of the present invention includes the following (A) and (B):

[0089] (A) a Cas3 protein, a polynucleotide encoding the protein, or an expression vector containing the polynucleotide; and (B) a cascade protein, a polynucleotide encoding the protein, or an expression vector containing the polynucleotide It may further comprise crRNA, a polynucleotide encoding said crRNA, or an expression vector containing said polynucleotide.

[0090] The components of the kit of the present invention may be in a form in which all or part of them are mixed together, or each component may be independent.

[0091] The kit of the present invention can be used in the fields of, for example, medicine, food, livestock, fisheries, industry, bioengineering, and life science research.

[0092] The kit of the present invention will be described below assuming that it is used for a pharmaceutical product (drug). When the kit is used in fields such as livestock farming, bioengineering, and life science research, the following description can be appropriately substituted based on common technical knowledge in the relevant fields.

[0093] Pharmaceuticals for editing DNA in animal cells, including humans, using the CRISPR-Cas3 system of the present invention can be prepared by conventional methods, more specifically, by compounding the molecules constituting the CRISPR-Cas3 system of the present invention with pharmaceutical additives, for example.

[0094] Here, the term "pharmaceutical additive" refers to a substance other than an active ingredient contained in a pharmaceutical. Pharmaceutical additives are substances contained in pharmaceuticals for purposes such as facilitating formulation, stabilizing quality, and enhancing usefulness. In one example, the pharmaceutical additives may be excipients, binders, disintegrants, lubricants, flow agents (anti-solidifying agents), colorants, capsule coatings, plasticizers, flavoring agents, sweeteners, flavoring agents, solvents, solubilizers, emulsifiers, suspending agents (adhesives), thickeners, pH adjusters (acidifiers, alkalinizers, buffers), wetting agents (solubilizers), antibacterial preservatives, chelating agents, suppository bases, ointment bases, hardening agents, softeners, medical water, propellants, stabilizers, and preservatives. These pharmaceutical additives can be readily selected by those skilled in the art based on the intended dosage form and administration route, as well as standard pharmaceutical practice.

[0095] Furthermore, the pharmaceutical composition for editing DNA in animal cells using the CRISPR-Cas3 system of the present invention may contain an additional active ingredient. The additional active ingredient is not particularly limited and can be appropriately designed by a person skilled in the art.

[0096] Specific examples of the active ingredients and pharmaceutical additives described above can be found, for example, in standards established by the US Food and Drug Administration (FDA), the European Medicines Agency (EMA), the Ministry of Health, Labour and Welfare of Japan, and the like.

[0097] Methods for delivering pharmaceuticals to desired cells include, for example, methods using viral vectors (adenoviral vectors, adeno-associated viral vectors, lentiviral vectors, Sendai viral vectors, etc.) that target the cells, or antibodies that specifically recognize the cells. Pharmaceuticals can take any dosage form depending on the purpose. Furthermore, the pharmaceuticals are prescribed appropriately by a doctor or medical professional.

[0098] The kit of the present invention preferably further comprises instructions for use. [Example]

[0099] The present invention will be described in more detail below with reference to examples, but the present invention is not limited to the following examples.

[0100] A. Establishing the CRISPR-Cas3 system in eukaryotic cells Materials and Methods [1] Preparation of a reporter vector containing the target sequence The target sequences were a sequence derived from the human CCR5 gene (sequence number 19) and an E. coli CRISPR spacer sequence (sequence number 22).

[0101] To insert the target sequence into the vector, we prepared a synthetic polynucleotide (SEQ ID NO: 20) containing a target sequence (SEQ ID NO: 19) derived from the human CCR5 gene, and a synthetic polynucleotide (SEQ ID NO: 21) containing a sequence complementary to the target sequence (SEQ ID NO: 19). Similarly, we prepared a synthetic polynucleotide (SEQ ID NO: 23) containing a target sequence (SEQ ID NO: 22) derived from the spacer sequence of E. coli CRISPR, and a synthetic polynucleotide (SEQ ID NO: 24) containing a sequence complementary to the target sequence (SEQ ID NO: 22). All of the above synthetic polynucleotides were obtained from Hokkaido System Science Co., Ltd.

[0102] The above polynucleotides were inserted into a reporter vector by the method described in [Sakuma T et al. (2013) Efficient TALEN construction and evaluation methods for human cell and animal applications, Genes to Cells, Vol. 18 (Issue 4), pp. 315-326]. The process is outlined below. First, polynucleotides having complementary sequences (polynucleotides of SEQ ID NO: 20 and SEQ ID NO: 21; polynucleotides of SEQ ID NO: 23 and SEQ ID NO: 24) were heated to 95°C for 5 minutes, then cooled to room temperature to allow hybridization. A block incubator (BI-515A, Astec Co., Ltd.) was used for this process. Next, the hybridized polynucleotides that formed a double-stranded structure were inserted into a substrate vector to create a reporter vector.

[0103] The sequences of the reporter vectors prepared are shown in SEQ ID NO: 31 (reporter vector containing a target sequence derived from the human CCR5 gene) and SEQ ID NO: 32 (reporter vector containing a target sequence derived from the spacer sequence of CRISPR in E. coli). The structure of the reporter vector is shown in Figure 4(d).

[0104] [2] Construction of Cse1 (Cas8), Cse2 (Cas11), Cas5, Cas6, Cas7, and crRNA expression vectors [Insert amplification and preparation] For the polynucleotides encoding Cse1 (Cas8), Cse2 (Cas11), Cas5, Cas6, and Cas7 (SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, and SEQ ID NO: 6, respectively), we first contracted GenScript to manufacture polynucleotides in the order SEQ ID NO: 2-SEQ ID NO: 3-SEQ ID NO: 6-SEQ ID NO: 4-SEQ ID NO: 5 (i.e., polynucleotides in which the sequences encoding each protein were linked in the order Cse1 (Cas8)-Cse2 (Cas11)-Cas7-Cas5-Cas6). The sequences encoding each protein, Cse1 (Cas8)-Cse2 (Cas11)-Cas7-Cas5-Cas6, were linked by the 2A peptide (amino acid sequence: GSGATNFSLLKQAGDVEENPGP (SEQ ID NO: 58)).

[0105] The nucleotide sequences encoding the 2A peptide differed slightly depending on the Cas protein junction, as follows: Sequence between Cse1 (Cas8) and Cse2 (Cas11): GGAAGCGGGAGCAACCAACTTCAGCCTGCTGAAGCAGGCCGGCGATGTGGAGGAGAATCCAGGCCCC (SEQ ID NO: 59); Sequence between Cse2 (Cas11) and Cas7: GGCTCCGGCGCCACCAATTTTTCTCTGCTGAAGCAGGCAGGCGATGTGGAGGAGAACCCAGGACCT (SEQ ID NO: 60); Sequence between Cas7 and Cas5: GGATCTGGAGCCACCAATTTCAGCCTGCTGAAGCAAGCAGGCGACGTGGAAGAAAACCCAGGACCA (SEQ ID NO: 61); Sequence between Cas5 and Cas6: GGATCTGGGGCTACTAATTTTTCTCTGCTGAAGCAAGCCGGCGACGTGGAAGAGAATCCAGGACCG (SEQ ID NO: 62).

[0106] Next, each polynucleotide was amplified under the PCR conditions (primers and time course) shown in the table below. PCR was performed using a 2720 Thermal cycler (applied biosystems).

[0107] [Table 1]

[0108] As a polynucleotide having a base sequence for expressing crRNA, a polynucleotide having the complementary sequence shown below was obtained. 1. Polynucleotide for expressing crRNA corresponding to a sequence derived from the human CCR5 gene (SEQ ID NOs: 25 and 26, obtained from Hokkaido System Science Co., Ltd.) 2. Polynucleotides for expressing crRNA corresponding to the spacer sequence of E. coli CRISPR (SEQ ID NOs: 27 and 28, obtained from Hokkaido System Science Co., Ltd.) 3. Polynucleotides for expressing crRNA corresponding to sequences derived from the human EMX1 gene (SEQ ID NOs: 29 and 30, obtained from FASMAC Co., Ltd.).

[0109] [Ligation and transformation] The base plasmid used was pPB-CAG-EBNXN (provided by the Sanger Center). 1.6 μg of the base plasmid was mixed with 1 μl of restriction enzymes BglII (New England Biolabs) and 0.5 μl of XhoI (New England Biolabs) in NEB buffer and reacted at 37°C for 2 hours. The digested base plasmid was purified using a Gel extraction kit (Qiagen).

[0110] The base plasmid and the insert prepared in this manner were ligated using the Gibson Assembly system at a base plasmid to insert ratio of 1:1, according to the Gibson Assembly system protocol (50°C for 25 minutes, total reaction volume: 8 μL).

[0111] Next, transformation was carried out by a conventional method using 6 μL of the obtained plasmid solution (ligation reaction solution) and competent cells (prepared by Takeda Laboratory).

[0112] The plasmid vector was then purified from the transformed E. coli by alkaline prep. Briefly, the plasmid vector was recovered using a QIAprep Spin Miniprep Kit (Qiagen), purified by ethanol precipitation, and then adjusted to a concentration of 1 μg / μL in TE buffer.

[0113] The structures of each plasmid vector are shown in Figure 4(a) to (c). The base sequences of the pre-crRNA expression vectors are shown in SEQ ID NO: 33 (expression vector expressing crRNA corresponding to a sequence derived from the human CCR5 gene), SEQ ID NO: 34 (expression vector expressing crRNA corresponding to a spacer sequence of E. coli CRISPR), and SEQ ID NO: 35 (expression vector expressing crRNA corresponding to a sequence derived from the human EMX1 gene).

[0114] [3] Construction of Cas3 expression vector A polynucleotide (SEQ ID NO: 1) encoding Cas3 and having a modified nucleotide sequence was obtained from Genscript, Inc. Specifically, a pUC57 vector incorporating the above polynucleotide was obtained from Genscript, Inc.

[0115] The vector was cleaved with the restriction enzyme NotI. The ends of the fragment were then blunted using 2 U of Klenow Fragment (Takara Bio) and 1 μL of 2.5 mM dNTP Mixture (Takara Bio). The fragment was then purified using gel extraction (Qiagen). The purified fragment was further cleaved with the restriction enzyme XhoI and purified using gel extraction (Qiagen).

[0116] The purified fragment was ligated using a base plasmid (pTL2-CAG-IRES-NEO vector, developed by the Takeda Laboratory) and a ligation kit (Mighty Mix, Takara Bio). After ligation, transformation and purification were performed using the same procedures as in [2]. The recovered plasmid vector was adjusted to a concentration of 1 μg / μL in TE buffer.

[0117] [4] Construction of a plasmid vector containing BPNLS Cas3, Cse1 (Cas8), Cse2 (Cas11), Cas5, Cas6, and Cas7 expression vectors were constructed with BPNLS linked to the 5' and 3' ends (see Figure 7).

[0118] We commissioned Thermo Fisher Scientific to create inserts for each Cas protein group containing BPNLS at both ends. The specific sequences of the inserts are (AGATCTTAATACGACTCACTATAGGGAGAGCCGCCACCATGGCC: SEQ ID NO: 56)-(any one of SEQ ID NOs: 7 to 12)-(TAATATCCTCGAG: SEQ ID NO: 57). SEQ ID NO: 56 is the sequence with a BgIII cleavage site. SEQ ID NO: 57 is the sequence with an XhoI cleavage site.

[0119] The pMK vector containing the above sequence was digested with the restriction enzymes BgIII and XhoI and purified using Gel Extraction (Qiagen). The purified fragment was ligated using a base plasmid (pPB-CAG-EBNXN, provided by the Sanger Center) and a ligation kit (Mighty Mix, Takara Bio). Transformation and purification were then carried out using the same procedures as in [2]. The recovered plasmid vector was adjusted to a concentration of 1 μg / μL in TE buffer.

[0120] [5] Construction of a plasmid vector containing Cascade (2A) We constructed an expression vector containing Cse1 (Cas8), Cse2 (Cas11), Cas7, Cas5, and Cas6 linked in this order. More specifically, we constructed an expression vector containing the following sequences: (NLS-Cse1 (Cas8): SEQ ID NO: 2)-2A-(NLS-Cse2 (Cas11): SEQ ID NO: 3)-2A-(NLS-Cas7: SEQ ID NO: 6)-2A-(NLS-Cas5: SEQ ID NO: 4)-2A-(NLS-Cas6: SEQ ID NO: 5) (see Figure 8). The amino acid sequence of the NLS is PKKKRKV (SEQ ID NO: 52), and the nucleotide sequence is CCCAAGAAGAAGCGGAAGGTG (SEQ ID NO: 53). The amino acid sequence of the 2A peptide is GSGATNFSLLKQAGDVEENPGP (SEQ ID NO: 58) (the corresponding nucleotide sequences are SEQ ID NOs: 59 to 62, respectively).

[0121] The polypeptide having the above-mentioned nucleotide sequence was obtained from GenScript. The pUC57 vector incorporating the above sequence was cleaved with the restriction enzyme EcoRI-HF and purified using Gel Extraction (Qiagen). The purified fragment was ligated using a base plasmid (pTL2-CAG-IRES-Puro vector, developed by the Takeda Laboratory) and a ligation kit (Mighty Mix, Takara Bio). Transformation and purification were then carried out using the same procedures as in [2]. The recovered plasmid vector was adjusted to a concentration of 1 μg / μL in TE buffer.

[0122] Example A-1 Cas3, Cse1 (Cas8), Cse2 (Cas11), Cas5, Cas6, and Cas7, each with a modified base sequence and a nuclear localization signal added, were expressed in HEK (human embryonic kidney) 293T cells along with crRNA, and their activity in cleaving the target sequence of exogenous DNA was evaluated.

[0123] Prior to transfection, HEK293T cells were cultured in a 10-cm dish in EF medium (GIBCO) at 37°C in a 5% CO atmosphere. The density of HEK293T cells in EF medium was 3 x 10 4 Prepared at 100 μL / 100 μL.

[0124] In addition, 100 ng of the above reporter vector; 200 ng each of Cas3 plasmid, Cse1 (Cas8) plasmid, Cse2 (Cas11) plasmid, Cas5 plasmid, Cas6 plasmid, Cas7 plasmid, and crRNA plasmid; 60 ng of pRL-TK vector (capable of expressing Renilla luciferase, Promega); and 300 ng of pBluescriptII KS(+) vector (Agilent Technologies) were mixed in 25 μL of Opti-MEM (Thermo Fisher Scientific). The conditions using a reporter vector with a CCR5-derived target sequence as the reporter vector correspond to 1 in Figure 1, and the conditions using a reporter vector with an E. coli CRISPR spacer sequence correspond to 10 in Figure 1.

[0125] 1.5 μL of Lipofectamine 2000 (Thermo Fisher Scientific) was mixed with 25 μL of OptiMEM (Thermo Fisher Scientific) and incubated at room temperature for 5 minutes. The plasmid + OptiMEM mixture was then mixed with the Lipofectamine 2000 + OptiMEM mixture and incubated at room temperature for 20 minutes. The resulting mixture was mixed with 1 mL of the above EF medium containing HEK293T cells and seeded into a 96-well plate (12 wells total, one well for each vector combination).

[0126] After culturing for 24 hours at 37°C under a 5% CO2 atmosphere, a dual-luciferase assay was performed according to the protocol of the Dual-Glo Luciferase assay system (Promega). 3 LB 960 (Berthold Technologies) was used.

[0127] As a control experiment, a similar experiment was carried out under the following conditions. 1. Instead of any one of the Cas3 plasmid, Cse1 (Cas8) plasmid, Cse2 (Cas11) plasmid, Cas5 plasmid, Cas6 plasmid, or Cas7 plasmid, an equal amount of pBluescriptII KS(+) vector (Agilent Technologies) was mixed and expressed (Figure 1, 2 to 7). 2. Instead of the crRNA plasmid used in the above procedure, a plasmid expressing a crRNA that is not complementary to the target sequence was mixed. That is, for the target sequence derived from the CCR5 gene, a plasmid expressing a crRNA corresponding to the spacer sequence of E. coli CRISPR was mixed (8 in Figure 1). When targeting the spacer sequence of E. coli CRISPR, a plasmid expressing a crRNA corresponding to the sequence derived from the CCR5 gene was mixed and expressed (11 in Figure 1). 3. As negative controls, only the reporter vector having the target sequence derived from CCR5 (9 in Figure 1) and only the reporter vector having the spacer sequence of E. coli CRISPR (12 in Figure 1) were expressed.

[0128] (result) The results of the dual-luciferase assay are shown in the graph at the top of Figure 1, and the experimental conditions are shown in the table at the bottom. In Figure 1(b), "CCR5-target" and "spacer-target" represent the target sequence derived from CCR5 and the spacer sequence of E. coli CRISPR, respectively. Furthermore, "CCR5-crRNA" and "spacer-crRNA" represent the sequences complementary to the CCR5-target and spacer-target, respectively.

[0129] In Figure 1, the system incorporating all of the Cas3, Cse1 (Cas8), Cse2 (Cas11), Cas5, Cas6, and Cas7 plasmids, along with a crRNA plasmid complementary to the target sequence, exhibited higher cleavage activity than the other systems (compare 1 with 2 to 8, and 10 with 11, respectively). Therefore, it was demonstrated that the expression vector according to one embodiment of the present invention can be used to express Cas3, Cse1 (Cas8), Cse2 (Cas11), Cas5, Cas6, and Cas7 in human cells.

[0130] Furthermore, it was suggested that introducing the above expression vector into human cells forms a complex of Cas3, Cascade, and crRNA in the human cells, cleaving the target sequence.

[0131] Furthermore, comparing 8 and 9, and 11 and 12 in Figure 1, the cleavage activity in the system expressing a crRNA that was not complementary to the target sequence was at a level equivalent to that of the negative control, suggesting that the CRISPR-Cas3 system of the present invention can specifically cleave sequences complementary to the crRNA in mammalian cells.

[0132] Example A-2 Using a method similar to that described in Example A-1, an experiment was conducted to evaluate whether the Type I CRISPR-Cas system can cleave endogenous DNA in human cells.

[0133] Specifically, we modified the base sequences of Cas3, Cse1 (Cas8), Cse2 (Cas11), Cas5, Cas6, and Cas7 with nuclear localization signals added, and expressed pre-crRNA in human cells, and evaluated whether the sequence of the endogenous CCR5 gene in the cells was cleaved.

[0134] The same HEK239T cells as in Example A-1 were cultured at 1 × 10 5 The cells were seeded at a density of 100 cells / well onto a 24-well plate and cultured for 24 hours.

[0135] 1 μg of Cas3 plasmid, 1.3 μg of Cse1 (Cas8) plasmid, 1.3 μg of Cse2 (Cas11) plasmid, 1.1 μg of Cas5 plasmid, 0.8 μg of Cas6 plasmid, 0.3 μg of Cas7 plasmid, and 1 μg of crRNA plasmid were mixed in 50 μL of Opti-MEM (Thermo Fisher Scientific). Next, a mixture of 5 μL of Lipofectamine® 2000 (Thermo Fisher Scientific), 50 μL of Opti-MEM (Thermo Fisher Scientific), and 1 mL of EF medium was added to the DNA mixture. Then, 1 mL of the resulting mixture was added to the 24-well plate.

[0136] After 24 hours of culture at 37°C in a 5% CO atmosphere, the medium was replaced with 1 mL of EF medium. 48 hours after transfection (24 hours after medium replacement), the cells were harvested and diluted to 1 × 10 4 The concentration was adjusted to 5 μL.

[0137] The cells were heated at 95° C. for 10 minutes. Next, 10 mg of proteinase K was added, and the mixture was incubated at 55° C. for 70 minutes. The mixture was further heated at 95° C. for 10 minutes and used as a template for PCR.

[0138] 10 μL of the above template was amplified by 35 cycles of two-step PCR. The PCR primers used were those having the sequences of SEQ ID NOs: 47 and 48. KOD FX (Toyobo Co., Ltd.) was used as the DNA polymerase, and the two-step PCR procedure followed the protocol attached to the KOD FX. The PCR-amplified product was purified using a QIAquick PCR Purification Kit (QIAGEN). The specific procedure followed the protocol attached to the kit.

[0139] dA was added to the 3' end of the purified DNA using rTaq DNA polymerase (Toyobo). The purified DNA was electrophoresed in a 2% agarose gel, and a band of approximately 500-700 bp was excised. DNA was then extracted and purified from the excised gel using a Gel Extraction Kit (QIAGEN). Next, TA cloning was performed using pGEM-T easy vector systems (Promega) to clone the DNA. Finally, the cloned DNA was extracted using the alkaline prep method and analyzed by Sanger sequencing. Analysis was performed using the BigDye® Terminator v3.1 Cycle Sequencing Kit (Thermo Fisher Scientific) and an Applied Biosystems 3730 DNA Analyzer (Thermo Fisher Scientific).

[0140] An overview of the endogenous CCR5 gene sequence targeted by the CRISPR-Cas system in this example will be described with reference to Figure 2. In Figure 2, exons are written in uppercase letters and introns are written in lowercase letters.

[0141] In this example, we targeted a sequence within the CCR5 gene located in the (P)21 region of the short arm of chromosome 3 (Figure 2; SEQ ID NO: 46 shows the full-length CCR5 nucleotide sequence). Specifically, we targeted a sequence within exon 3 of the CCR5 gene. As a control, a Cas9 target sequence was also placed at approximately the same position. That is, the entire underlined sequence is the target sequence of the Type I CRISPR-Cas system (AAG and the following 32 bases), and the double-underlined sequence is the target sequence of Cas9 (CGG and the preceding 20 bases). We designed the crRNA sequence to enable guidance to the target sequence of the Type I CRISPR-Cas system (AAG and the following 32 bases).

[0142] (result) As a result of the above experiment, clone 1, clone 2, clone 3, and clone 4 were obtained, each with a 401-bp deletion, a 341-bp deletion, a 268-bp deletion, and a 344-bp deletion, respectively, compared to the original base sequence (Figures 3A-D). This demonstrates that the CRISPR-Cas3 system of the present invention can delete endogenous DNA in human cells. This suggests that the CRISPR-Cas system can be used to edit DNA in human cells.

[0143] In this example, clones with deleted base pairs were observed, which supports the idea that the CRISPR-Cas3 system of the present invention generates DNA cleavage at multiple sites.

[0144] The CRISPR-Cas3 system of the present invention deleted several hundred base pairs (268-401 bp) of DNA, which was more extensive than the deletions achieved by the CRISPR-Cas system using Cas9 (which usually cuts only one site on the DNA).

[0145] Example A-3 Using a method similar to that of Example A-1, an experiment was conducted to evaluate whether the CRISPR-Cas3 system can cleave endogenous DNA in human cells.

[0146] Specifically, we modified the base sequences of Cas3, Cse1 (Cas8), Cse2 (Cas11), Cas5, Cas6, and Cas7 with nuclear localization signals added, and expressed pre-crRNA in human cells, and evaluated whether the sequence of the endogenous EMX1 gene in the cells was cleaved.

[0147] The same HEK293T cells as in Example A-1 were cultured at 1 × 10 5 The cells were seeded onto a 24-well plate at a density of 100 cells / well and cultured for 24 hours.

[0148] 500 ng of Cas3 plasmid, 500 ng of Cse1 (Cas8) plasmid, 1 μg of Cse2 (Cas11) plasmid, 1 μg of Cas5 plasmid, 1 μg of Cas6 plasmid, 3 μg of Cas7 plasmid, and 500 μg of crRNA plasmid were mixed in 50 μL of Opti-MEM (Thermo Fisher Scientific). 4 μL of Lipofectamine® 2000 (Thermo Fisher Scientific) and 50 μL of Opti-MEM (Thermo Fisher Scientific) were added to the mixture and mixed. The resulting mixture was incubated at room temperature for 20 minutes and then added to the HEK293T cells.

[0149] The structure of the expression vector for the Cas proteins used in Example A-3 is shown in Figure 7. As shown in Figure 7, the expression vector contains a sequence encoding the Cas proteins flanked by bipartite NLSs (BPNLS) on both sides (see Suzuki K et al. (2016) In vivo genome editing via CRISPR / Cas9 mediated homology-independent targeted integration, Nature, Vol. 540 (Issue 7631), pp. 144-149). The amino acid sequence of BPNLS is KRTADGSEFESPKKKRKVE (SEQ ID NO: 54), and the nucleotide sequence is AAGCGGACTGCTGATGGCAGTGAATTTGAGTCCCCAAAGAAGAAGAGAAAGGTGGAA (SEQ ID NO: 55).

[0150] The HEK293T cells were cultured at 37°C under a 5% CO atmosphere for 24 hours, and then the medium was replaced with 1 mL of EF medium (1 mL per well). 48 hours after transfection (24 hours after medium replacement), the cells were harvested and diluted to 1 × 10 in PBS. 4 The concentration was adjusted to 5 μL.

[0151] The cells were heated at 95° C. for 10 minutes. Next, 10 mg of proteinase K was added, and the mixture was incubated at 55° C. for 70 minutes. The mixture was further heated at 95° C. for 10 minutes and used as a template for PCR.

[0152] 10 μL of the above template was amplified by 40 cycles of 3-step PCR. The PCR primers used were those with the sequences of SEQ ID NOs: 50 and 51. Hotstartaq (QIAGEN) was used as the DNA polymerase, and the 3-step PCR procedure followed the protocol provided with the Hotstartaq. The PCR-amplified product was electrophoresed in a 2% agarose gel, and a band of approximately 900 to 1100 bp was excised. DNA was then extracted and purified from the excised gel using a Gel Extraction Kit (QIAGEN). The specific procedure followed the protocol provided with the kit.

[0153] Next, the DNA was cloned using pGEM-T easy vector systems (Promega) via transfection cloning. Finally, the cloned DNA was extracted using the alkaline prep method and analyzed by Sanger sequencing. The BigDye® Terminator v3.1 Cycle Sequencing Kit (Thermo Fisher Scientific) and an Applied Biosystems 3730 DNA Analyzer (Thermo Fisher Scientific) were used.

[0154] An outline of the endogenous EMX1 gene sequence targeted by the CRISPR-Cas3 system in Example A-3 is described based on Figure 5. In Figure 5, exons are written in uppercase letters and introns are written in lowercase letters.

[0155] In Example A-3, we targeted a sequence within the EMX1 gene located in the short arm (P)13 region of chromosome 2 (Figure 5; SEQ ID NO: 49 shows the full-length nucleotide sequence of EMX1). Specifically, we targeted a sequence within exon 3 of the EMX1 gene. As a control, a Cas9 target sequence was placed at approximately the same position. That is, the underlined sequence further upstream is the target sequence of the Type I CRISPR-Cas system (AAG and the following 32 bases), and the underlined sequence further downstream is the target sequence of Cas9 (TGG and the preceding 20 bases). The crRNA sequence used in Example A-3 was designed to enable guidance to the target sequence of the CRISPR-Cas3 system (AAG and the following 32 bases).

[0156] (result) As a result of the above experiment, clone 1, which had two deletions of 513 bp and 363 bp compared to the original base sequence, and clone 2, which had a deletion of 694 bp, were obtained (Figures 6A and 6B). These experimental results also demonstrated that the CRISPR-Cas3 system of the present invention can delete endogenous DNA in human cells. In other words, it was suggested that the CRISPR-Cas3 system can edit DNA in human cells.

[0157] Furthermore, similar to Example A-2, the double-stranded DNA was cleaved at two or more sites, and several hundred base pairs of DNA were deleted. Therefore, the results of Example A-3 more strongly support the suggestion obtained from Example A-2.

[0158] Example A-4 The CRISPR-Cas3 system, in which the base sequence was modified and a base sequence encoding a cascade protein was further linked, was expressed in HEK293T cells, and its activity in cleaving the target sequence of exogenous DNA was evaluated.

[0159] In Example A-4, 100 ng of reporter vector, 200 ng each of Cas3 plasmid, Cascade (2A) plasmid, and crRNA plasmid, 60 ng of pRL-TK vector (capable of expressing Renilla luciferase, Promega), and 300 ng of pBluescriptII KS(+) vector (Agilent Technologies) were mixed in 25 μL of Opti-MEM (Thermo Fisher Scientific). The condition using a reporter vector with a CCR5-derived target sequence as the reporter vector corresponds to 1 in (b) of Figure 9, and the condition using a reporter vector with an E. coli CRISPR spacer sequence corresponds to 6 in (b) of Figure 9.

[0160] The reporter vectors used were the two types of reporter vectors constructed in [1] of [Preparation Example] (i.e., vectors with the structure shown in Figure 4(d)).The Cascade (2A) plasmid used was the expression vector constructed in [4] of [Preparation Example] (i.e., vectors with the structure shown in Figure 8).

[0161] A dual-luciferase assay was carried out in the same manner as in Example A-1, except that the above expression vector was used.

[0162] As a control experiment, a similar experiment was conducted under the following conditions. 1. Instead of either the Cas3 plasmid or the Cascade (2A) plasmid, an equal amount of pBluscriptII KS(+) vector (Agilent Technologies) was mixed and expressed (2 and 3 in Figure 9). 2. Instead of the crRNA plasmid used in the above procedure, a plasmid expressing a crRNA that is not complementary to the target sequence was mixed. That is, for the target sequence derived from the CCR5 gene, a plasmid expressing a crRNA corresponding to the spacer sequence of the E. coli CRISPR was mixed (4 in Figure 9). When targeting the spacer sequence of the E. coli CRISPR, a plasmid expressing a gRNA corresponding to the sequence derived from the CCR5 gene was mixed and expressed (7 in Figure 9). 3. As negative controls, only the reporter vector having the target sequence derived from CCR5 (5 in Figure 9) and only the reporter vector having the spacer sequence of E. coli CRISPR (8 in Figure 9) were expressed.

[0163] (result) The results of the dual-luciferase assay are shown in the graph at the top of Figure 9, and the experimental conditions are shown in the table at the bottom of Figure 9. In Figure 9, "CCR5-target" and "spacer-target" represent the target sequence derived from CCR5 and the spacer sequence of E. coli CRISPR, respectively. Furthermore, "CCR5-crRNA" and "spacer-crRNA" represent the sequences complementary to the CCR5-target and the spacer-target, respectively.

[0164] As shown in Figure 9, the system incorporating both the Cas3 and Cascade (2A) plasmids and the crRNA plasmid complementary to the target sequence exhibited significantly higher cleavage activity than the other systems (compare 1 with 2 to 5, and 6 with 7 to 8). This suggests that the CRISPR-Cas system according to one embodiment of the present invention can specifically cleave sequences complementary to the crRNA in mammalian cells, even in systems in which the nucleotide sequences encoding the Cascade proteins are linked and expressed.

[0165] B. Verification of factors affecting genome editing using the CRISPR-Cas3 system in eukaryotic cells Materials and Methods [1] Cas gene and crRNA structure Cas3 and Cascade component genes (Cse1, Cse2, Cas5, Cas6, and Cas7) derived from Escherichia coli K-12 were engineered with bpNLSs at their 5' and 3' ends, and then cloned into mammalian cells by gene synthesis after codon optimization. These genes were subcloned downstream of the CAG promoter in the pPB-CAG.EBNXN plasmid, a gift from the Sanger Institute. Cas3 mutants, including H74A (dead nickase; dn), K320N (dead helicase; dh), and the double mutant S483A and T485A (dead helicase ver. 2; dh2), were generated by self-ligation of PCR products from PrimeSTAR MAX. For the crRNA expression plasmid, a crRNA sequence containing two BbsI restriction enzyme sites at the spacer position under the U6 promoter was synthesized. All crRNA expression plasmids were constructed by inserting a 32-base pair double-stranded oligo of the target sequence into the BbsI restriction enzyme site.

[0166] The Cas9-sgRNA expression plasmid, pX330-U6-Chimeric_BB-CBh-hSpCas9, was obtained from Addgene. To design the gRNA, we used the CRISPR web tool, CRISPR design tool, and / or CRISPRdirect, which predict unique target sites in the human genome. The target sequence was cloned into the sgRNA scaffold of pX330 according to the protocol of the Feng Zhang Laboratory.

[0167] The SSA reporter plasmid containing two BsaI restriction enzyme sites was a gift from Professor Taku Yamamoto of Hiroshima University. The target sequence of the genomic region was inserted into the BsaI site. The Renilla luciferase vector pRL-TK (Promega) was obtained. All plasmids were prepared by midiprep or maxiprep using the PureLink HiPure Plasmid Purification Kit (Thermo Fisher).

[0168] [2] Evaluation of DNA cleavage activity in HEK293T cells To detect DNA cleavage activity in mammalian cells, an SSA assay was performed as in Example A. HEK293T cells were cultured in high-glucose Dulbecco's modified Eagle's medium (Thermo Fisher) supplemented with 10% fetal bovine serum at 37°C under 5% CO2. 0.5 × 10 4 Cells were seeded into wells of a 96-well plate. 24 hours later, HEK293T cells were transfected with 100 ng each of Cas3, Cse1, Cse2, Cas7, Cas5, Cas6, and crRNA expression plasmids, 100 ng of an SSA reporter vector, and 60 ng of a Renilla luciferase vector using lipofectamine 2000 and OptiMEM (Life Technologies) according to a slightly modified protocol. 24 hours after transfection, dual luciferase assays were performed using the Dual-Glo luciferase assay system (Promega) according to the protocol.

[0169] [3] Detection of indels in HEK293T cells 2.5x10 4Twenty-four hours after seeding cells into wells of a 24-well plate, HEK293T cells were transfected with 250 ng of Cas3, Cse1, Cse2, Cas7, Cas5, Cas6, and crRNA expression plasmids (each 250 ng) using lipofectamine 2000 and OptiMEM (Life Technologies) according to a slightly modified protocol. Two days after transfection, total DNA was extracted from harvested cells using a Tissue XS kit (Takara Bio) according to the protocol. Target loci were amplified using Gflex (Takara Bio) or Quick Taq HS DyeMix (TOYOBO) and subjected to agarose gel electrophoresis. To detect small insertion / deletion mutations in PCR products, the SURVEYOR Mutation Detection Kit (Integrated DNA Technologies) was used according to the protocol. For TA cloning, the pCR4Blunt-TOPO plasmid vector (Life Technologies) was used according to the protocol. For sequence analysis, the BigDye Terminator Cycle Sequencing Kit and the ABI PRISM 3130 Genetic Analyzer (Life Technologies) were used.

[0170] To detect rare variants, DNA libraries were prepared from PCR amplification products using the TruSeq Nano DNA Library Prep Kit (Illumina), and amplicon sequencing was performed on a MiSeq (2 x 150 bp) according to Macrogen's standard procedures. Raw reads from each sample were mapped to hg38 of the human genome using BWA-MEM. Coverage data were visualized using Integrative Genomics Viewer (IGV), and histograms were extracted for the target regions.

[0171] To detect SNP-KI (Snip knock-in) in mammalian cells, reporter HEK293T cells carrying mCherry-P2A-EGFP c321C>G were kindly provided by Professor Shinichiro Nakata. Reporter cells were cultured in 1 μg / ml puromycin. 500 ng of donor plasmid or single-stranded DNA was co-transfected with CRISPR-Cas3 as described above. Five days after transfection, all cells were harvested and subjected to FACS analysis using an AriaIIIu (BD). GFP-positive cells were sorted, and total DNA was extracted as described above. Genome-wide SNP exchange was detected by PCR amplification using HiDi DNA polymerase (myPOLS Biotec).

[0172] [4] Detection of off-target site candidates Type IE CRISPR off-target candidates were identified in the human genome hg38 using GGGenome with two different approaches. The PAM candidate sequences were selected as AAG, ATG, AGG, GAG, TAG, and AAC, as previously reported (Leenay, RT, et al. Mol. Cell 62, 137-147 (2016); Jung, et al. Mol. Cell. 2017; Jung et al. Cell 170, 35-47 (2017)). Because positions with multiples of six have been reported to be unrecognized as target sites (Kunne et al. Molecular Cell 63, 1-13 (2016)), the first approach excluded these positions and selected sequences with the fewest mismatches within the 32 base pairs of the target sequence. The second approach identified regions with perfect matches to the 5' end of the PAM side of the target sequence and listed them in descending order of their mismatch potential.

[0173] [5] Deep sequencing for off-target analysis For whole-genome sequencing, genomic DNA was extracted from transfected HEK293T cells and sheared using a Covaris sonicator. DNA libraries were prepared using the TruSeq DNA PCR-Free LT Library Prep Kit (Illumina), and genome sequencing was performed using a HiSeq X (2 × 150 bp) instrument according to Takara Bio's standard procedures. Raw reads from each sample were mapped to hg38 of the human genome using BWA-MEM and cleaned using the Trimmomatic program. Discordant read pairs and split reads were filtered out using samtools and Lumpy-sv, respectively. To detect only large deletions on the same chromosome, read pairs mapped to different chromosomes were removed using BadMateFilter in the Genome Analysis Toolkit program. The total number of discordant read pairs or split reads in each 100-kb region was counted using Bedtools, and the error rate relative to the negative control was calculated. To enrich for potential off-target regions prior to sequencing, SureSelectXT custom DNA probes were designed by SureDesign under moderately stringent conditions and manufactured by Agilent Technologies. The target regions were selected as follows: Probes near the target region covered 800 kb upstream and 200 kb downstream of the PAM. Near the CRISPR-Cas3 off-target region, probes covered 9 kb upstream and 1 kb downstream of the PAM candidate. Near the CRISPR-Cas9 off-target region, probes covered 1 kb upstream and 1 kb downstream of the PAM. After DNA library preparation using the SureSelectXT reagent kit and custom probe kit, genome sequencing was performed on a Hiseq 2500 (2 × 150 bp) according to Takara Bio's standard procedures. Discordant read pairs and split reads on the same chromosome were excluded using the methods described above.The total number of discordant read pairs or split reads in each 10 kb region was counted using Bedtools, and the error rate compared with the negative control was calculated.

[0174] [Example B-1] Effect of the type of crRNA and nuclear localization signal on DNA cleavage activity In Example A, we successfully performed genome editing in eukaryotic cells using the CRISPR-Cas3 system, which coincidentally contained a pre-crRNA (LRSR; leader sequence-repeat sequence-spacer sequence-repeat sequence) as the crRNA. The inventors suspected that the lack of success in genome editing in eukaryotic cells using the CRISPR-Cas3 system over the years may be due to the use of mature crRNA as the crRNA. Therefore, in addition to pre-crRNA (LRSR), we also prepared pre-crRNA (RSR; repeat sequence-spacer sequence-repeat sequence) and mature crRNA (5' handle sequence-spacer sequence-3' handle sequence) as crRNAs, and examined the genome editing efficiency using the reporter system described in Example A (Figures 10A and 10B). The nucleotide sequences of pre-crRNA (LRSR), pre-crRNA (RSR), and mature crRNA are shown in SEQ ID NOs: 63, 64, and 65, respectively.

[0175] As a result, the CRISPR-Cas3 system using mature crRNA did not show any target DNA cleavage activity. Surprisingly, however, when pre-crRNA (LRSR, RSR) was used, very high target DNA cleavage activity was observed. This result with the CRISPR-Cas3 system contrasts with the CRISPR-Cas9 system, which shows high DNA cleavage activity when using mature crRNA. This fact also suggests that one of the main reasons why the CRISPR-Cas3 system has not been successful in genome editing in eukaryotic cells so far is that it has relied on mature crRNA.

[0176] We also tested the use of the SV40 nuclear localization signal and the bipartite nuclear localization signal as nuclear localization signals added to Cas3 (Figure 11). As a result, higher target DNA cleavage activity was observed when the bipartite nuclear localization signal was used.

[0177] Therefore, in the subsequent experiments, we used pre-crRNA (LRSR) as the crRNA and the bipartite nuclear localization signal as the nuclear localization signal.

[0178] [Example B-2] Effect of PAM sequence on DNA cleavage activity To confirm the target specificity of the CRISPR-Cas3 system, we investigated the effect of various PAM sequences on DNA cleavage activity (Figure 12). In the SSA assay, different PAM sequences resulted in variable DNA cleavage activity. 5'-AAG PAM showed the highest activity, with AGG, GAG, TAC, ATG, and TAG also showing notable activity.

[0179] [Example B-3] Effect of mismatch between crRNA and spacer sequence on DNA cleavage activity Previous studies of the crystal structure of the E. coli Cascade showed that the crRNA and spacer DNA form a 5-nt heteroduplex, which is caused by the disruption of base pairing at every sixth position by the thumb element of the Cas7 effector (Fig. 13). We assessed the effect of mismatches between the crRNA and spacer sequences on DNA cleavage activity (Fig. 1g). With the exception of the unrecognized base (position 6), any single mismatch in the seed region (positions 1–8) dramatically reduced cleavage activity.

[0180] [Example B-4] Verification of the necessity of each domain of Cas3 for DNA cleavage activity In vitro characterization of the catalytic properties of the Cas3 protein revealed that the N-terminal HD nuclease domain cleaves a single-stranded region of the DNA substrate, followed by the C-terminal SF2 helicase domain, which unwinds the target DNA in a 3'-to-5' direction in an ATP-dependent manner. We constructed three Cas3 mutants—a mutant with an HD domain H74A (dnCas3), a mutant with a K320N mutation in SF2 domain motif 1 (dhCas3), and a double mutant with S483A / T485A mutation in SF2 domain motif 3 (dh2Cas3)—to verify whether the Cas3 domain is required for DNA cleavage (Figure 14). All three Cas3 mutants completely lost DNA cleavage activity, indicating that Cas3 can cleave target DNA via the HD nuclease and SF2 helicase domains.

[0181] [Example B-5] Verification of DNA cleavage activity in various types of CRISPR-Cas3 systems Type 1 CRISPR-Cas3 systems are highly diverse (seven types, Type 1 A to Type 1 G). While the previous example examined the DNA cleavage activity of Type IE CRISPR-Cas3 systems in eukaryotic cells, this example examined the DNA cleavage activity of other Type 1 CRISPR-Cas3 systems (Type IF and Type IG). Specifically, we codon-optimized and cloned Type IF Cas3 and Cas5-7 from Shewanella putrefaciens and Type IG Cas5-8 from Pyrococcus furiosus (Figure 15). Although the strength of DNA cleavage activity varied, these Type 1 CRISPR-Cas3 systems also demonstrated DNA cleavage activity in SSA assays using 293T cells.

[0182] [Example B-6] Verification of mutations introduced into endogenous genes using the CRISPR-Cas3 system We verified the mutations introduced into endogenous genes by the CRISPR-Cas3 system using the type IE system. The EMX1 and CCR5 genes were selected as target genes, and pre-crRNA (LRSR) plasmids were prepared. 293T cells were lipofected with pre-crRNA and plasmids encoding six Cas(3,5-8,11) effectors. CRISPR-Cas3-mediated deletions of hundreds to thousands of base pairs were primarily located upstream of the 5' PAM of the spacer sequence in the target region (Figure 16). Microhomology of 5-10 base pairs was observed at the repaired junction, which may be due to annealing of complementary strands via the annealing-dependent repair pathway. Genome editing in the EMX1 and CCR5 regions was not observed with the mature crRNA plasmid.

[0183] To further characterize Cas3-mediated genome editing by TA cloning and Sanger sequencing of PCR products, 96 TA clones were selected and compared with the wild-type EMX1 sequence (Figure 17). Of the 49 clones in which insertion was confirmed, 24 clones contained deletions of a minimum of 596 base pairs, a maximum of 1447 base pairs, and an average of 985 base pairs (46.3% efficiency). Half of the clones (n = 12) contained large deletions including the PAM and spacer sequences, while the other half contained deletions upstream of the PAM.

[0184] Further characterization of Cas3 was performed by next-generation sequencing of PCR amplification products using primer sets covering broader regions, including the 3.8 kb region of the EMX1 gene and the 9.7 kb region of CCR5. We also tested multiple PAM sites (AAG, ATG, and TTT) for type IE CRISPR targeting. Amplicon sequencing revealed significantly reduced coverage across broad genomic regions upstream of the PAM sites, with AAG at 38.2% and ATG at 56.4%, compared with the 86.4% coverage achieved with TTT and 86.4% coverage achieved with Cas9 targeting EMX1. The reduced coverage was also observed when targeting the CCR5 region. In contrast, Cas9 induced small insertions and deletions (indels) at the target site, whereas Cas3 induced no small indel mutations at the PAM or target site. These results suggest that the CRISPR-Cas3 system can induce deletions across broad regions upstream of the target site in human cells.

[0185] Given the limitations of PCR analysis, such as amplification of less than 10 kb and a strong bias favoring shorter PCR fragments, we utilized microarray-based capture sequencing of over 1000 kb surrounding the targeted EMX1 and CCR5 loci (Fig. 18A, B). We observed deletions of up to 24 kb at the EMX1 locus and up to 43 kb at the CCR5 locus, but 90% of EMX1 and 95% of CCR5 mutations were less than 10 kb. These results suggest that the CRISPR-Cas3 system can possess potent nuclease and helicase activities even in eukaryotic genomes.

[0186] As shown with the CRISPR-Cas9 system, the possibility of inducing unwanted off-target mutations in non-target genomic regions is a major concern, especially for clinical applications. However, no significant off-target effects were observed with the CRISPR-Cas3 system. [Industrial Applicability]

[0187] The CRISPR-Cas3 system of the present invention can edit the DNA of eukaryotic cells, and therefore can be widely used in fields where genome editing is required, such as medicine, agriculture, forestry, fisheries, industry, life sciences, bioengineering, and gene therapy.

Claims

1. A method for producing eukaryotic cells (excluding cells in a state in which they exist in the human body, human germ cells, and human embryonic cells) or animals (excluding humans) with edited DNA, the method comprising introducing into the eukaryotic cells (excluding cells in a state in which they exist in the human body, human germ cells, and human embryonic cells) or animals (excluding humans) a CRISPR-Cas3 system capable of editing DNA in the eukaryotic cells or animals, wherein the CRISPR-Cas3 system comprises any of the following (A) to (C) (excluding when the CRISPR-Cas3 system is Type I-D): (A) a Cas3 protein, a polynucleotide encoding the protein, or an expression vector containing the polynucleotide; (B) a Cascade protein, a polynucleotide encoding the protein, or an expression vector containing the polynucleotide; and (C) a pre-crRNA, a polynucleotide encoding the crRNA, or an expression vector containing the polynucleotide

2. A method for producing a DNA-edited plant, comprising introducing into the plant a CRISPR-Cas3 system capable of editing DNA within the plant, wherein the CRISPR-Cas3 system comprises the following (A) to (C) (excluding when the CRISPR-Cas3 system is Type I-D): (A) a Cas3 protein, a polynucleotide encoding the protein, or an expression vector containing the polynucleotide; (B) a Cascade protein, a polynucleotide encoding the protein, or an expression vector containing the polynucleotide; and (C) a pre-crRNA, a polynucleotide encoding the crRNA, or an expression vector containing the polynucleotide

3. The method according to claim 1 or 2, comprising a step of introducing a CRISPR-Cas3 system into a eukaryotic cell, an animal, or a plant, and then cleaving crRNA with a protein constituting a cascade protein.

4. The method according to any one of claims 1 to 3, wherein a nuclear localization signal is added to the Cas3 protein and / or the Cascade protein.

5. The method of claim 4, wherein the nuclear localization signal is a bipartite nuclear localization signal.

6. A kit for use in producing eukaryotic cells or animals with edited DNA, the kit comprising the following (A) to (C) that constitute a CRISPR-Cas3 system capable of editing DNA in eukaryotic cells or animals (excluding cases where the CRISPR-Cas3 system is Type I-D): (A) a Cas3 protein, a polynucleotide encoding the protein, or an expression vector containing the polynucleotide; (B) a Cascade protein, a polynucleotide encoding the protein, or an expression vector containing the polynucleotide; and (C) a pre-crRNA, a polynucleotide encoding the crRNA, or an expression vector containing the polynucleotide

7. A kit for use in producing a DNA-edited plant, the kit comprising the following (A) to (C) constituting a CRISPR-Cas3 system capable of editing DNA in a plant (excluding cases where the CRISPR-Cas3 system is Type I-D): (A) a Cas3 protein, a polynucleotide encoding the protein, or an expression vector containing the polynucleotide; (B) a Cascade protein, a polynucleotide encoding the protein, or an expression vector containing the polynucleotide; and (C) a pre-crRNA, a polynucleotide encoding the crRNA, or an expression vector containing the polynucleotide

8. The kit according to claim 6 or 7, wherein a nuclear localization signal is added to the Cas3 protein and / or the Cascade protein.

9. The kit of claim 8, wherein the nuclear localization signal is a bipartite nuclear localization signal.

Citation Information

Patent Citations

  • Genome engineering with type i crispr systems in eukaryotic cells

    WO2017066497A2

  • Target sequence specific alteration technology using nucleotide target recognition

    WO2019039417A1

  • Modified CASCADE ribonucleoproteins and their uses

    JP2015503535A

  • Methods and compositions for RNA-dependent transcriptional repression using crispr-related genes

    JP2017512481A