NOVEL MINIATURE CRISPR-Cas12

The novel small CRISPR-Cas12 protein, TD1, addresses the challenge of large molecular size in existing CRISPR-Cas systems by enabling efficient cell transfection and improved genome editing efficacy through its smaller size and specific DNA cleavage capabilities.

WO2025094873A1PCT designated stage expired Publication Date: 2025-05-08ISHINO YOSHIZUMI +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/038286
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-30
Filing Date
2024-10-28
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Existing CRISPR-Cas systems, such as Cas9 and Cas12a, have large molecular sizes, which hinder efficient transfection into cells and limit the effectiveness of genome editing techniques.

Method used

Development of a novel small CRISPR-Cas12 protein, TD1, which has approximately 500 amino acid residues, significantly smaller than Cas9 and Cas12a, and is capable of forming a complex with crRNA to target and cleave DNA.

Benefits of technology

The smaller size of TD1 enhances its efficiency in transfecting into cells and improves the efficacy of genome editing, while maintaining the ability to specifically recognize and cleave target DNA sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024038286_08052025_PF_FP_ABST
    Figure JP2024038286_08052025_PF_FP_ABST
Patent Text Reader

Abstract

According to the present invention, novel CRISPR-Cas is searched, and a Cas protein obtained from the same is provided. The protein is one among the following (a), (b), and (c): (a) a protein composed of the amino acid sequence of SEQ ID NO: 3; (b) a protein composed of an amino acid sequence that shares an identity of at least 62% and less than 100% with the amino acid sequence of SEQ ID NO: 3, forming a complex with a crRNA, and capable of binding to DNA having a sequence complementary to a spacer sequence in the crRNA; and (c) a protein composed of a sequence in which 1-20 amino acids are mutated in the amino acid sequence of SEQ ID NO: 3, forming a complex with a crRNA, and capable of binding to DNA having a sequence complementary to a spacer sequence in the crRNA. Moreover, provided are: a nucleic acid encoding said protein; a recombinant vector; and a transformed cell. Also provided is a kit or method for genome editing.
Need to check novelty before this filing date? Find Prior Art

Description

Novel small CRISPR-Cas12

[0001] The invention relates to a novel miniature CRISPR-Cas12 that can be used as a tool for genetic engineering, such as genome editing.

[0002] CRISPR (clustered regularly interspaced short palindromic repeat) is a DNA sequence found in the genomes of prokaryotic bacteria and archaea, where repeats of several tens of base pairs (repeats) are arranged at regular intervals. It has been shown to function in prokaryotic adaptive immunity (Non-Patent Document 1). CRISPRs are also closely associated with CRISPR-associated (cas) gene clusters, which encode nucleases and helicases. Typical CRISPRs contain three elements: cas genes, a leader sequence, and a repeat / spacer sequence. However, CRISPR-Cas are highly diverse and are divided into class 1 and class 2 based on the type of Cas protein and its mechanism of action. Class 1 includes types I, III, and IV, while class 2 includes types II, V, and VI. It has also been proposed that each type can be further subdivided into several subtypes (Non-Patent Document 2). The CRISPR-Cas system is a system in which CRISPR RNA (crRNA), transcribed from DNA sequences inserted between repeats, forms a complex with Cas nuclease, which guides the Cas nuclease to DNA sites with sequences homologous to the crRNA, resulting in double-stranded cleavage of the DNA at those sites. Practical genome editing techniques utilizing this system have been developed and are widespread worldwide (Non-Patent Document 3). The most widely used method utilizes Cas9, derived from the bacterium Streptococcus pyrogenes, a type II CRISPR-Cas system. Subsequently, CRISPR-Cas derived from the bacterium Francisella novicida came into use. The nuclease protein was originally named Cpf1, but is now called Cas12a in line with the classification of CRISPR-Cas systems.

[0003] The CRISPR-Cas system has not only been used as a genome editing technology, but also has the ability to freely target proteins to targeted locations on the genome, leading to the development of a variety of applications, including transcriptional control of specific genes, site-specific epigenetic conversion, site-specific imaging, and nucleic acid detection methods (Non-Patent Documents 4 and 5).

[0004] Barrangou R. et al., Science, Vol. 315, No. 6, pp. 1709-1712, 2007; Makarova K. S. et al., Nat Rev Microbiol., Vol. 18, No. 11, pp. 67-83, 2019; Jinek M. et al., Science, Vol. 337, pp. 810-821, 2012; Knott G. J. et al., Science, Vol. 361, pp. 6405, pp. 866-869, 2018; Mazhar A. et al., Nature Communications, Vol. 9, pp. 1911, 2018

[0005] CRISPR-Cas9 is a versatile system that forms the core of genome editing, and CRISPR-Cas12 has recently attracted attention. However, these Cas proteins have a problem of large molecular size and low cellular introduction efficiency. Therefore, to improve genome editing technology, it is essential to discover new Cas proteins with small molecular size. Furthermore, the shorter the crRNA used to guide the Cas effector with DNA cleavage activity to the target DNA site, the more advantageous it is. The objective of the present invention is to provide a novel, small Cas protein.

[0006] We previously analyzed metagenomic data from marine samples and identified numerous CRISPR arrays. Furthermore, we identified an ORF near the CRISPR array that is predicted to have an amino acid sequence similar to RuvC, the active domain of Cas12. These ORFs consist of approximately 500 amino acid residues, less than half the size of Cas9 or Cas12. Aiming for application to genome editing, we performed functional and structural analysis of a small Cas12 protein, designated TD1. TD1 was coexpressed with a CRISPR array in Escherichia coli and purified as an RNA-bound complex. RNA-seq analysis revealed that the RNA bound to TD1 was the processed crRNA transcribed from the CRISPR array. To determine the PAM sequence required for TD1-targeted double-stranded DNA cleavage, we constructed a plasmid library containing random nucleotide sequences at the 5' end of the protospacer. The TD1-crRNA complex and the plasmid library were mixed, and the TD1-cleaved plasmid was analyzed by next-generation sequencing to determine the TD1 PAM sequence. We also confirmed that the target DNA cleavage activity of TD1 depends on conserved acidic residues within the putative RuvC domain. The TD1-crRNA-target DNA ternary complex was reconstituted and its structure determined at 2.9 Å resolution by single-particle analysis using cryo-electron microscopy. Mutation analysis based on structural information was performed to elucidate the mechanism of target DNA cleavage by TD1. The present invention was completed based on these findings. The gist of the present invention is as follows: (1) A protein of any of the following (a), (b), or (c): (a) a protein consisting of the amino acid sequence of SEQ ID NO: 3; (b) a protein consisting of an amino acid sequence with 62% or more but less than 100% identity to the amino acid sequence of SEQ ID NO: 3, which forms a complex with crRNA and can bind to DNA having a sequence complementary to the spacer sequence in crRNA; or (c) a protein consisting of the amino acid sequence of SEQ ID NO: 3 with 1 to 20 amino acid mutations, which forms a complex with crRNA and can bind to DNA having a sequence complementary to the spacer sequence in crRNA. (2) The protein described in (1) above, which has nuclease activity.(3) The protein described in (1) that does not have nuclease activity. (4) A nucleic acid encoding the protein described in (1). (5) An expression vector containing DNA encoding the protein described in (1). (6) A cell transformed with the expression vector described in (5). (7) A method for producing any of the following proteins (a), (b), or (c), comprising culturing the cell described in (6). (a) A protein consisting of the amino acid sequence of SEQ ID NO: 3. (b) A protein consisting of an amino acid sequence having 62% or more but less than 100% identity with the amino acid sequence of SEQ ID NO: 3, which can form a complex with crRNA and bind to DNA having a sequence complementary to the spacer sequence in crRNA. (c) A protein consisting of the amino acid sequence of SEQ ID NO: 3 with 1 to 20 amino acids mutated, which can form a complex with crRNA and bind to DNA having a sequence complementary to the spacer sequence in crRNA. (8) A method for cleaving DNA having a sequence complementary to the spacer sequence in crRNA using the protein and crRNA described in (1). (9) A method for genome editing using the protein and crRNA described in (1). (10) A genome editing kit comprising at least one selected from the group consisting of the protein of (1), mRNA encoding the protein of (1), and an expression vector comprising DNA encoding the protein of (1). (11) A crRNA having a repeat sequence comprising any of the nucleotide sequences of SEQ ID NOs: 4 to 8. (12) The protein of (1), wherein the protein consists of an amino acid sequence homologous to the amino acid sequence of SEQ ID NO: 3, and wherein acidic or basic amino acids at positions corresponding to Asp156, Arg191, and Arg193 in the amino acid sequence of SEQ ID NO: 3 form hydrogen bonds with bases in the PAM sequence of the target DNA. (13) The crRNA of (11), wherein the spacer forms an R-loop with the target DNA.

[0007] The present invention provides a novel CRISPR-Cas and a method for producing the same. The novel CRISPR-Cas belongs to Class 2, Type V in currently known classification systems and functions as a single peptide as an effector involved in cleaving target DNA, providing a practical genome editing tool. Furthermore, the size of this effector peptide is significantly smaller than those currently used in genome editing, providing a tool that is easy to handle and has higher introduction efficiency. This specification incorporates the contents of the specification and / or drawings of Japanese Patent Application No. 2023-185496, from which the present application claims priority.

[0008] Environmental metagenomic analysis procedure. Seawater collected from a fixed point in Sendai Bay was filtered stepwise, and DNA was prepared from the microorganisms retained on a 0.22 μm filter and subjected to NGS sequence analysis. Discovery of CRISPR-Cas regions from metagenomic sequences. The resulting sequence data was assembled to obtain as many contiguous sequences as possible. A database was then created to search for CRISPR-like repeat sequences, resulting in the discovery of the gene region shown in the figure (the base sequence is shown in SEQ ID NO: 1). The arrows indicate the region translated into amino acid sequence and its translation direction. Co-production of TD1 with surrounding regions containing the CRISPR array. To identify the crRNA that binds to the TD1 protein, we assumed the RNA transcribed from the CRISPR region to be crRNA and added it to TD1. The downstream CRISPR region was inserted into an E. coli expression vector, and the constructed plasmid was introduced into E. coli cells for co-expression. When E. coli cell extract was subjected to Ni-NTA affinity chromatography, heparin affinity chromatography, and gel filtration using the His-Tag attached to TD1, TD1 was eluted as a single peak, and the peak fraction was A 260 Absorption of A 280The peaks were larger than the 280 nm peak, indicating the presence of nucleic acids. Identification of nucleic acids bound to TD1. Nucleic acid extraction was performed on the 280 nm and 260 nm peaks in Figure 3, revealing two specific bands. Subsequent treatment with DNase and RNase only eliminated the bands after RNase treatment, indicating that the nucleic acid contained with the TD1 protein was RNA. The results of RNA sequencing analysis are mapped onto the plasmid. The horizontal axis represents each position, and the vertical axis represents the number of reads from RNA sequencing. The highest number of reads was at the CRISPR array position, 60 nt, indicating that the crRNA was transcribed and processed from the CRISPR array. This suggests that the crRNA functions as a guide RNA for TD1. PAM sequence recognized by TD1. PAM sequences are necessary to distinguish between the genome and foreign nucleic acids. After cleavage of the target DNA by TD1, the gap was filled, an adapter was added, and analysis was performed by NGS. The results showed that TD1 specifically recognizes the C at position -2 of the 5'-HCN-3' sequence. Identifying the cleavage site by TD1. Based on the sequence analysis shown in Figure 6, we estimated the cleavage site of TD1. NTS cleaved 19 (20) nt downstream of the PAM, and TS cleaved 22 (21) nt downstream of the PAM, forming a 5'-overhanging end. PAM sequence specificity. Using one of the recognized PAMs, 5'-GATCA-3', we created sequences in which the 1st, 2nd, 3rd, and 4th bases from the 3' end were changed to four different bases, and the cleavage efficiency was compared. The horizontal axis of the bar graph represents each base, and the vertical axis represents the percentage of DNA cleaved. No significant differences were observed at positions -4 and -1. At position -3, TD1 preferred A, T, and C over G, resulting in cleavage only at position -2, where C was the only cleavage site. This suggests that TD1 specifically cleaves at position -2, C, which is consistent with the NGS results. TD1 cleavage temperature. The TD1-crRNA complex and a plasmid containing C at position -2 of the PAM were incubated at temperatures ranging from 4°C to 50°C. The results showed that the target DNA was significantly cleaved at temperatures between 37 and 44°C, suggesting the possibility of application to human cells. The TD1 mutant (a mutant in which aspartic acid at position 308 was converted to alanine) was created.The TD1-crRNA complex, which had been mutated to eliminate DNA cleavage activity, was purified and mixed with target dsDNA, incubated, and reconstituted by gel filtration chromatography. TD1, crRNA, and DNA were detected in the [1] peak, and structural analysis was performed using cryo-EM. Three-dimensional image of the TD1-crRNA complex and target DNA with elimination of DNA cleavage activity. The peak fraction [1] in Figure 10 was subjected to cryo-EM observation. From the resulting 5,447 electron microscope images, 1,248,508 particles considered to be of interest were selected, and similar structures were classified to create a 2D average image. Based on this, an initial 3D structure was constructed, and a refined 3D image was reconstructed using a dataset believed to be TD1 as a reference. A map with a resolution of 2.9 Å was obtained. Three-dimensional image of crRNA. In the refined 3D image in Figure 11, the crRNA repeats formed a pseudoknot structure, and the spacer formed an R-loop with the target DNA. Identification of amino acids capable of recognizing PAM. Hydrogen bonds were found between NTS-2 C and Asp156, TS-2 G and Arg191, and TS-3 A and Arg193, respectively, demonstrating base-specific recognition by TD1. Functional structural analysis revealed that TD1 is a Cas protein that functions as a monomer, suggesting its potential application as a genome editing tool with a smaller molecular size than existing Cas proteins.

[0009] Hereinafter, an embodiment of the present invention will be described.

[0010] The present invention provides any one of the following proteins (a), (b), or (c): (a) a protein consisting of the amino acid sequence of SEQ ID NO: 3; (b) a protein consisting of an amino acid sequence with at least 62% but less than 100% identity to the amino acid sequence of SEQ ID NO: 3, which can form a complex with crRNA and bind to DNA having a sequence complementary to a spacer sequence in crRNA; or (c) a protein consisting of the amino acid sequence of SEQ ID NO: 3 with 1 to 20 amino acid mutations, which can form a complex with crRNA and bind to DNA having a sequence complementary to a spacer sequence in crRNA. The amino acid sequence of SEQ ID NO: 3 is closest to that of ARI, a Cas protein discovered by the present inventors (Patent Application No. 2023-148437), with 61.7% identity. Sequence identity can be calculated using analytical software such as MAFFT.

[0011] The protein consisting of the amino acid sequence of SEQ ID NO: 3 (protein (a) above) preferably has nuclease activity. This activity involves forming a complex with crRNA, which is produced by processing pre-crRNA produced by transcription of the CRISPR region. Depending on the sequence of the crRNA, the protein targets a DNA strand having a complementary sequence to the crRNA, guiding a polypeptide consisting of the amino acid sequence of SEQ ID NO: 3 to that location and cleaving both strands. However, the protein of the present invention does not necessarily have nuclease activity. The protein of the present invention can form a complex with crRNA and bind to DNA (target DNA) having a sequence complementary to the spacer sequence in the crRNA. The crRNA that forms a complex with the protein of the present invention preferably has a repeat sequence comprising a CRISPR repeat sequence (double underlined) in the sequence of SEQ ID NO: 1 (the sequence of the gene region encoding the novel CRISPR-Cas) minus the four 5' bases (GATT). The sequences in which the four 5' bases (GATT) are minus the CRISPR repeat sequence in SEQ ID NO: 1 are shown in SEQ ID NOs: 4 to 8. CRISPR repeat sequences are generally considered diverse, and the CRISPR repeat sequences present in SEQ ID NO: 1 contain one, two, or seven base differences compared to the initial CRISPR repeat sequence. CRISPR repeat sequences differing in these bases may have equivalent functions. The length of the repeat sequence is preferably 23-55 bp, more preferably 24-38 bp, and more preferably 29-30 bp. Furthermore, the crRNA complexed with the protein of the present invention preferably contains a spacer that forms an R-loop with the target DNA. The length of the spacer sequence is preferably 26-51 bp, more preferably 30-45 bp. For the Cas effector protein to act on the target DNA, Cas and crRNA form a complex, and Cas is guided to DNA with a sequence homologous to the spacer RNA portion of the crRNA. Cas recognizes the PAM sequence present there, thereby exerting its function of cleaving the double-stranded DNA.Structural analysis of the Cas-crRNA-DNA complex of the present invention identified the amino acid sequence by which the Cas of the present invention recognizes PAM and confirmed that the crRNA cleaves the double-stranded target DNA and forms a hybrid strand with one of the strands (R-loop formation). The repeat sequence in the crRNA used in this study was a 36-mer, 5'-GATTAAGGCCCTTGTGTAGTGGGGTGTAACTACAAC-3'. However, the four 5'-nucleotides (GATT) were not observed in the resulting structure, likely indicating that they were cleaved during analysis or that the structure was unstable. Therefore, it is believed that a 32-mer crRNA, 5'-AAGGCCCTTGTGTAGTGGGGTGTAACTACAAC-3' (SEQ ID NO: 4), is sufficient for the Cas-crRNA of the present invention to exert its intended function. The binding of the complex to the target DNA can be confirmed directly by cleaving the target DNA (if the protein of the present invention has nuclease activity), labeling the target DNA (if the protein of the present invention does not have nuclease activity but is labeled), using various molecular biological experimental techniques depending on the labeling method, or by crystal structure analysis, etc. In the Examples described below, it was confirmed that Asp156, Arg191, and Arg193 in the protein (TD1) consisting of the amino acid sequence of SEQ ID NO: 3 form hydrogen bonds with bases in the PAM sequence of the target DNA, based on the construction of a three-dimensional model obtained by electron microscopy and the cleavage activity of the TD1 mutant.

[0012] The protein of the present invention and crRNA form a complex, and the crRNA guides the protein of the present invention to the target DNA strand. However, the tracrRNA required by the currently widely used Cas9 is not necessarily required. In other words, this property can be achieved with a simpler molecular structure. Furthermore, a PAM (protospacer adjacent motif) sequence of 5'-HC-3' (H=A / T / C) that distinguishes self DNA from target DNA enables specific cleavage of the target sequence.

[0013] To cleave a target DNA strand or guide the protein of the present invention to a target DNA site, crRNA can be artificially synthesized. In genome editing, crRNA is generally referred to as guide RNA, and the polypeptide that cleaves the target DNA is often referred to as effector. The guide RNA only needs to form base pairs with the target DNA when complexed with the effector. There is no limit to its length, but it is typically about 100 nucleotides. To achieve the desired activity, the effector and guide RNA can be prepared separately and mixed at the time of use, or they can be prepared as a complex and used in the reaction.

[0014] The protein (b) above has an amino acid sequence that shares 62% or more but less than 100% identity with the amino acid sequence of SEQ ID NO: 3, forms a complex with crRNA, and can bind to DNA having a sequence complementary to the spacer sequence in the crRNA. The identity between the amino acid sequence of the protein (b) above and the amino acid sequence of SEQ ID NO: 3 is 62% or more but less than 100%, and is preferably 65% ​​or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, or 95% or more but less than 100%. Sequence identity can be calculated using analytical software such as MAFFT.

[0015] The protein (c) above consists of an amino acid sequence of SEQ ID NO: 3 with 1 to 20 amino acids mutated, forms a complex with crRNA, and can bind to DNA having a sequence complementary to the spacer sequence in crRNA.

[0016] The protein (c) above preferably has an amino acid sequence in which one or more amino acids have been deleted, substituted, or added in the amino acid sequence of SEQ ID NO: 3. The total number of deleted, substituted, or added amino acids is one or more, and the specific range is usually 1 to 20, 1 to 15, 1 to 10, preferably 1 to 5, and more preferably 1 to 2 for deletions; usually 1 to 20, 1 to 15, 1 to 10, preferably 1 to 5, and more preferably 1 to 2 for substitutions; and usually 1 to 20, 1 to 15, 1 to 10, preferably 1 to 5, and more preferably 1 to 2 for additions. The proteins (b) and (c) above have an amino acid sequence homologous to the amino acid sequence of SEQ ID NO: 3, and in that amino acid sequence, acidic or basic amino acids at positions corresponding to Asp156, Arg191, and Arg193 in the amino acid sequence of SEQ ID NO: 3 can form hydrogen bonds with bases in the PAM sequence of the target DNA.

[0017] The protein of the present invention may be either a glycosylated or non-glycosylated protein. The type, position, etc. of the glycosylated sugar chain added to a protein will vary depending on the type of host cell used to produce the protein, but glycosylated proteins include proteins obtained using any host cell.

[0018] The present invention also provides nucleic acids encoding any of the proteins (a), (b), or (c) above. The nucleic acid may be any of single-stranded DNA, single-stranded RNA (e.g., mRNA), a single-stranded polynucleotide consisting of a mixture of DNA and RNA, double-stranded DNA, double-stranded RNA, a DNA-RNA hybrid polynucleotide, or a double-stranded polynucleotide consisting of two polynucleotides consisting of a mixture of DNA and RNA. Any of the proteins (a), (b), or (c) above can be produced using DNA encoding any of the proteins (a), (b), or (c). Furthermore, RNA (mRNA) encoding any of the proteins (a), (b), or (c) above can be used for genome editing.

[0019] The DNA sequence encoding the protein consisting of the amino acid sequence of SEQ ID NO: 3 is shown in SEQ ID NO: 2. DNA having the sequence of SEQ ID NO: 2 can be produced by known artificial gene synthesis methods. For example, a series of overlapping oligonucleotides are synthesized and annealed to form a double-stranded DNA fragment containing nicks in both strands. Repairing these nicks with DNA ligase yields an artificial gene having the desired sequence. RNA (mRNA) having the sequence of SEQ ID NO: 2 can be produced by preparing a plasmid DNA encoding the sequence of SEQ ID NO: 2, amplifying and purifying the template DNA by PCR, and then performing a transcription reaction using RNA polymerase.

[0020] DNA encoding the above proteins (b) and (c) can be obtained by artificially introducing mutations into DNA having the sequence of SEQ ID NO: 2 using known methods such as site-directed mutagenesis. Mutations can be introduced using, for example, a mutagenesis kit, such as Mutan-K (manufactured by TAKARA), Mutan-G (manufactured by TAKARA), TAKARA's LA PCR in vitro Mutagenesis series kit, or a kit based on a similar principle commercially available from various companies. The proteins can also be obtained by preparing oligoDNA containing mutations at the desired positions and chemically synthesizing the entire gene region using the method described above.

[0021] The protein of the present invention can be produced, for example, by constructing a recombinant vector by inserting DNA encoding the protein of the present invention into an expression vector, transforming host cells with this recombinant vector, culturing the transformed cells, and purifying the protein produced by the transformed cells.

[0022] When constructing a recombinant vector, a DNA fragment of an appropriate length containing the coding region of a protein of interest is first prepared. Nucleotides may be substituted in the nucleotide sequence of the coding region of the protein of interest so that the codons are optimal for expression in a host cell.

[0023] Next, the DNA fragment is inserted downstream of a promoter of an appropriate expression vector to prepare a recombinant vector. The DNA fragment must be incorporated into the vector so that its function can be exerted, and the vector may contain, in addition to the promoter, cis elements such as enhancers, splicing signals, poly(A) addition signals, selection markers (e.g., dihydrofolate reductase gene, ampicillin resistance gene, neomycin resistance gene), ribosome binding sequences (SD sequences), etc.

[0024] The expression vector is not particularly limited as long as it is capable of autonomous replication in host cells, and examples of suitable vectors include plasmid vectors, phage vectors, and viral vectors. Examples of suitable plasmid vectors include E. coli-derived plasmids (e.g., pRSET, pBR322, pBR325, pUC118, pUC119, pUC18, and pUC19), Bacillus subtilis-derived plasmids (e.g., pUB110 and pTP5), and yeast-derived plasmids (e.g., YEp13, YEp24, and YCp50). Examples of suitable phage vectors include λ phage (e.g., Charon4A, Charon21A, EMBL3, EMBL4, λgt10, λgt11, and λZAP). Examples of suitable viral vectors include retroviruses, animal viruses such as vaccinia virus, adenovirus, and adeno-associated virus (AAV), and insect viruses such as baculovirus.

[0025] A recombinant vector in which DNA encoding the protein of the present invention is inserted into an expression vector can also be used for genome editing.

[0026] Transformed cells capable of producing the desired protein can be obtained by introducing a recombinant vector into a suitable host cell.

[0027] As the host cell, any of prokaryotic cells, yeast, animal cells, insect cells, plant cells, etc. may be used as long as it is capable of expressing DNA encoding the target protein. Also, animal individuals, plant individuals, silkworm bodies, etc. may be used.

[0028] When a bacterium is used as a host cell, for example, a bacterium belonging to the genus Escherichia such as Escherichia coli, the genus Bacillus such as Bacillus subtilis, the genus Pseudomonas such as Pseudomonas putida, or the genus Rhizobium such as Rhizobium meliloti can be used as the host cell. Specifically, Escherichia coli such as Escherichia coli BL21, Escherichia coli XL1-Blue, Escherichia coli XL2-Blue, Escherichia coli DH1, Escherichia coli K12, Escherichia coli JM109, and Escherichia coli HB101, and Bacillus subtilis such as Bacillus subtilis MI 114 and Bacillus subtilis 207-21 can be used as the host cell. In this case, the promoter is not particularly limited as long as it can be expressed in bacteria such as E. coli. Examples of the promoter include the trp promoter, the lac promoter, and the P L Promoter, P R Promoters derived from E. coli or phages, such as the promoter, can be used. Artificially designed and modified promoters, such as the tac promoter, lacT7 promoter, and letI promoter, can also be used.

[0029] The method for introducing a recombinant vector into bacteria is not particularly limited as long as it is a method that can introduce DNA into bacteria, and for example, a method using calcium ions, electroporation, etc. can be used.

[0030] When yeast is used as the host cell, Saccharomyces cerevisiae, Schizosaccharomyces pombe, Pichia pastoris, etc. can be used as the host cell. In this case, the promoter is not particularly limited as long as it can be expressed in yeast, and examples that can be used include the gal1 promoter, gal10 promoter, heat shock protein promoter, MFα1 promoter, PHO5 promoter, PGK promoter, GAP promoter, ADH promoter, and AOX1 promoter.

[0031] The method for introducing a recombinant vector into yeast is not particularly limited as long as it is a method that can introduce DNA into yeast, and for example, electroporation, spheroplast method, lithium acetate method, etc. can be used.

[0032] When animal cells are used as host cells, monkey cells COS-7, Vero, Chinese hamster ovary cells (CHO cells), mouse L cells, rat GH3, human FL cells, etc. can be used as host cells. In this case, the promoter is not particularly limited as long as it can be expressed in animal cells, and examples that can be used include the SRα promoter, SV40 promoter, LTR (Long Terminal Repeat) promoter, CMV promoter, and early gene promoter of human cytomegalovirus.

[0033] The method for introducing a recombinant vector into animal cells is not particularly limited as long as it is a method that can introduce DNA into animal cells, and for example, electroporation, calcium phosphate method, lipofection method, etc. can be used.

[0034] When insect cells are used as hosts, Spodoptera frugiperda ovarian cells, Trichoplusia ni ovarian cells, cultured cells derived from silkworm ovaries, etc. can be used as host cells. Examples of Spodoptera frugiperda ovarian cells that can be used include Sf9 and Sf21, Trichoplusia ni ovarian cells such as High 5 and BTI-TN-5B1-4 (Invitrogen), and cultured cells derived from silkworm ovaries such as Bombyx mori N4.

[0035] The method for introducing a recombinant vector into insect cells is not particularly limited as long as it allows DNA to be introduced into insect cells, and for example, the calcium phosphate method, lipofection, electroporation, etc. can be used.

[0036] The target protein can be produced by culturing transformed cells into which a recombinant vector incorporating DNA encoding the target protein has been introduced. The transformed cells can be cultured according to a conventional method used for culturing host cells.

[0037] As a medium for culturing transformed cells obtained using a microorganism such as Escherichia coli or yeast as a host, either a natural medium or a synthetic medium may be used as long as it contains a carbon source, a nitrogen source, inorganic salts, etc. that can be utilized by the microorganism and allows the transformed cells to be cultured efficiently.

[0038] Examples of carbon sources that can be used include carbohydrates such as glucose, fructose, sucrose, and starch, organic acids such as acetic acid and propionic acid, and alcohols such as ethanol and propanol. Examples of nitrogen sources that can be used include inorganic acids or ammonium salts of organic acids such as ammonia, ammonium chloride, ammonium sulfate, ammonium acetate, and ammonium phosphate, peptone, meat extract, yeast extract, corn steep liquor, and casein hydrolysate. Examples of inorganic salts that can be used include monopotassium phosphate, dipotassium phosphate, magnesium phosphate, magnesium sulfate, sodium chloride, ferrous sulfate, manganese sulfate, copper sulfate, and calcium carbonate.

[0039] Transformants obtained using microorganisms such as Escherichia coli or yeast as hosts can be cultured under aerobic conditions, such as shaking culture or aeration and agitation culture. The culture temperature is usually 25 to 37°C, the culture time is usually 12 to 48 hours, and the pH is usually maintained at 6 to 8 during the culture period. The pH can be adjusted using inorganic acids, organic acids, alkaline solutions, urea, calcium carbonate, ammonia, etc. During culture, antibiotics such as ampicillin and tetracycline may be added to the medium, if necessary.

[0040] When culturing a microorganism transformed with an expression vector using an inducible promoter, an inducer may be added to the medium as needed. For example, isopropyl-β-D-thiogalactopyranoside or the like may be added to the medium when culturing a microorganism transformed with an expression vector using the lac promoter, and indoleacrylic acid or the like may be added to the medium when culturing a microorganism transformed with an expression vector using the trp promoter.

[0041] As a medium for culturing transformed cells obtained using animal cells as hosts, commonly used RPMI1640 medium, Eagle's MEM medium, DMEM medium, Ham's F12 medium, Ham's F12K medium, or media containing these media supplemented with fetal bovine serum, etc., can be used. Transformed cells are usually cultured in an atmosphere of 5% CO 2 The incubation is carried out in the presence of β-actin at 37° C. for 3 to 10 days. During incubation, antibiotics such as kanamycin, penicillin, streptomycin, etc. may be added to the medium, if necessary.

[0042] As a medium for culturing transformed cells obtained using insect cells as hosts, commonly used media such as TNM-FH medium (Pharmingen), Sf-900 II SFM medium (Gibco BRL), ExCell400, and ExCell405 (JRH Biosciences) can be used. Transformed cells are typically cultured at 27°C for 3 to 10 days. During culture, antibiotics such as gentamicin may be added to the medium, if necessary.

[0043] The target protein may be expressed as a secreted protein or a fusion protein, such as β-galactosidase, protein A, the IgG-binding region of protein A, chloramphenicol acetyltransferase, poly(Arg), poly(Glu), protein G, maltose-binding protein, glutathione S-transferase, polyhistidine tether (His-tag), S-peptide, a DNA-binding protein domain, Tac antigen, thioredoxin, or green fluorescent protein.

[0044] The target protein can be obtained by collecting it from the culture of the transformed cells. Here, the "culture" includes any of the culture supernatant, cultured cells, cultured bacterial cells, and cell or bacterial lysate.

[0045] When the target protein accumulates intracellularly in the transformed cells, the cells are collected by centrifugation of the culture, washed, and then disrupted to extract the target protein. When the target protein is secreted extracellularly in the transformed cells, the culture supernatant is used as is, or the cells or bacterial bodies are removed from the culture supernatant by centrifugation or the like.

[0046] The obtained protein can be purified by solvent extraction, salting-out desalting with ammonium sulfate or the like, precipitation with an organic solvent, diethylaminoethyl (DEAE)-Sepharose, ion exchange chromatography, hydrophobic chromatography, gel filtration, affinity chromatography, or the like.

[0047] The protein of the present invention can also be produced based on its amino acid sequence by chemical synthesis methods such as the Fmoc method (fluorenylmethyloxycarbonyl method) and the tBoc method (t-butyloxycarbonyl method), using a commercially available peptide synthesizer.

[0048] The protein of the present invention forms a complex with crRNA (guide RNA), and the crRNA can guide the protein to a target DNA strand. This property can be utilized to enable genome editing. The present invention provides a method for genome editing using the protein (a), (b), or (c) and crRNA (guide RNA) described above.

[0049] If the protein of the present invention has nuclease activity, it can cleave the target DNA strand guided by the crRNA (guide RNA), thereby enabling DNA modification. The present invention also provides a method for cleaving DNA having a sequence complementary to the spacer sequence in the crRNA using the protein and crRNA described above in (a), (b), or (c). Gene targeting using CRISPR requires the coexpression of a target sequence-specific crRNA (guide RNA) and a nuclease (Cas protein) in cells. The protein of the present invention and crRNA (guide RNA) may be expressed from a single vector, or the protein of the present invention and crRNA (guide RNA) may be coexpressed by expressing them from separate vectors. DNA double-strand breaks induced by the protein of the present invention and crRNA (guide RNA) are repaired by either nonhomologous end joining (NHEJ) or homology-directed repair (HDR). In NHEJ, repeated cleavage at the target site frequently leads to repair errors, resulting in base insertions and deletions. This disrupts gene function (knockout). In addition, HDR repairs the damage using the homologous region of the unbroken chromosome as a template. This allows for the replacement of DNA sequences at specific sites on a chromosome or the insertion of foreign DNA into specific sites (knock-in). Specifically, a DNA fragment (donor DNA) containing the DNA to be replaced or inserted is prepared between sequences before and after the target site on the chromosome. By coexisting this donor DNA at the site of the DNA double-strand break, the donor DNA is inserted into the break site (knock-in). Genetic modification through genome editing makes it possible to elucidate gene function, improve breeds, treat diseases, and create disease model animals that mimic human genetic abnormalities.

[0050] If the protein of the present invention does not have nuclease activity, it can be used to regulate (activate or repress) the transcription of a specific gene. Furthermore, by labeling the protein of the present invention with a fluorescent dye, biotin, an enzyme, or the like, it can be used for site-specific imaging, nucleic acid detection, and the like. Furthermore, by linking an enzyme that modifies the epigenome to the protein of the present invention, site-specific epigenetic conversion becomes possible.

[0051] The present invention also provides a kit for genome editing. The kit of the present invention comprises at least one selected from the group consisting of the following (i), (ii), and (iii): (i) the protein (a), (b), or (c) above, (ii) mRNA encoding the protein (a), (b), or (c) above, and (iii) an expression vector comprising DNA encoding the protein (a), (b), or (c) above.

[0052] The above protein (a), (b) or (c) may or may not have nuclease activity, and may be labeled with a fluorescent dye, biotin, an enzyme or the like.

[0053] The mRNA encoding the protein (a), (b), or (c) above may be chemically modified with pseudouridine, 5-methylcytosine, 5-methoxyuridine, or the like, or may be codon-optimized, and may contain a cap structure, a poly-A tail, or the like.

[0054] For expression vectors containing DNA encoding the above proteins (a), (b), or (c), viral vectors such as adenoviruses and adeno-associated viruses (AAVs) or plasmid vectors can be used. In addition to promoters, vectors can contain cis elements such as enhancers, splicing signals, poly(A) addition signals, selection markers (e.g., dihydrofolate reductase genes, ampicillin resistance genes, neomycin resistance genes), ribosome binding sequences (SD sequences), and the like. The expression vectors may be linearized. The expression vectors may also contain repeat sequences of crRNA.

[0055] The kit of the present invention may further include a buffer, a crRNA as a negative control, a crRNA as a positive control, instructions for use, etc.

[0056] The present invention is described in detail below. [Example 1] Metagenomic Analysis. The inventors fractionated seawater collected from a fixed location in Sendai Bay, Tohoku, by filtering it through three filters with different pore sizes (8 μm, 1 μm, and 0.22 μm) to separate it based on cell size. DNA samples prepared from the bacterial / archaeal fraction (retained by the 0.22 μm filter) and the viral / phage fraction (passed through the 0.22 μm filter) were subjected to sequence analysis using a next-generation sequencer (NGS) (Figure 1). The analysis was performed using Illumina's Miseq. The resulting metagenomic sequence library was used for sequence assembly to construct as contiguous sequence data as possible. A search for CRISPR-like repeat sequences within this database revealed several putative CRISPRs. Detailed sequence comparison of candidate Cas proteins located nearby identified an open reading frame (ORF) encoding a type V Cas effector-like sequence (Figure 2). To reanalyze the metagenomic sequence surrounding the ORF, the original DNA prepared from seawater and subjected to NGS was used to amplify the region by PCR, and the sequence was confirmed by dideoxy sequencing, which is more accurate than NGS. To analyze the function of the polypeptide encoded by this gene, we constructed an expression plasmid by inserting the gene into the vector pET21a(+) (Merck) for expression in E. coli. This expression plasmid was then transformed into the E. coli host BL21(DE3 CodonPlus-RIL (Agilent) to produce the target polypeptide. However, purification was difficult, likely due to instability. Therefore, we constructed the pETDuet1-6His-TD1+array by inserting a neighboring CRISPR-like sequence into pETDuet1 (Merck) for co-expression. E. coli BL21(DE3) CodonPlus-RIL was transformed and cultured with this expression plasmid. As a result, the production efficiency of the target polypeptide (designated TD1) increased, and we were able to purify TD1 to a high purity, as described below (Figure 3).

[0057] Coexpression of the polypeptide encoded by the ORF and its surrounding region. The expression vector was designed so that the expressed polypeptide would have six histidines (His) at the amino terminus. Therefore, 6His-TD1 produced in E. coli was first affinity purified using Ni-NTA agarose (Qiagen). The polypeptide that specifically bound to Ni-NTA agarose was fractionated and further purified to a high purity by affinity chromatography using heparin (HiTrap Heparin HP, Cytiva). This fraction was concentrated by ultrafiltration and then loaded onto a Superdex 200 Increase 10 / 300 GL (Cytiva). The main peak was confirmed to contain 6His-TD1 (Figure 3). Furthermore, the UV absorption of the fraction eluted with the main peak by gel filtration was higher at 260 nm than at 280 nm (A 280 < A 260 ), it was predicted that this fraction contained nucleic acids (Figure 3).

[0058] Identification of Nucleic Acids Binding to TD1 To identify the nucleic acids contained in the purified fractions obtained by co-expressing the CRISPR region and TD1 protein, the components contained in the peak fractions obtained by gel filtration were separated into polypeptides and nucleic acids by phenol-chloroform-isoamyl alcohol (PCI) extraction, and the nucleic acids were recovered by ethanol precipitation. The nucleic acids were then treated with DNase or RNase and analyzed by denaturing PAGE (Figure 4). The results revealed that the nucleic acids contained in the complex were resistant to DNase degradation but were degraded by RNase degradation, and consisted of RNA approximately 50 strands long (Figure 4). The isolated RNA was subjected to sequence analysis (RNAseq) using NGS. Of all reads, 98% were derived from the pETDuet1-6His-TD1+ array, with significantly higher expression levels in the CRISPR array region (Figure 5). The most highly expressed region was found to match the CRISPR spacer and part of the repeat sequence. This suggests that the RNA bound to TD1 is crRNA, consisting of a repeat and a portion of the spacer derived from the CRISPR region. Therefore, it is thought that TD1 itself processes the crRNA precursor transcribed from the CRISPR region and cleaves it to produce crRNA of the appropriate length. In other words, it is thought that TD1 recognizes and binds to the stem-loop structure of the repeat, cleaving it at two points: upstream of the stem-loop and in the middle of the spacer.

[0059] Specific Cleavage Activity of Target DNA and Selection of PAMs We constructed an experimental system to examine target DNA cleavage using purified TD1-crRNA complexes. The target DNA was the 5'-most spacer sequence in the CRISPR array. The DNA substrate was constructed using pUC18. A short sequence called a Protospacer Adjacent Motif (PAM) is required for Cas effectors to recognize their target sequences. Because this sequence varies depending on the individual Cas effector, its sequence must also be determined for TD1. To achieve this, we prepared a PAM library (5'-NNNNN-3') containing a mixture of A, G, C, and T at each position of a five-base sequence and inserted it adjacent to the target DNA. The target sequence was a fragment of the human DNMT1 gene (an enzyme that methylates cytosines in genomic DNA). TD1 cleavage using this plasmid as a substrate confirmed that the circular plasmid was cleaved and linearized. Using this reaction system, we then determined the PAM sequence and cleavage site preferred by TD1. Because we did not know how double-stranded DNA cleavage occurred, we first blunted the ends of the cleavage reaction product, added an A, and then ligated an adapter with a T overhang. PCR amplification was then performed using primers based on the adapter sequence. Next, we sequenced the amplified product using NGS to examine the sequence at the PAM position. We observed specificity, with a C at the second position from the 3' end of 5'-NNNNN-3' (Figure 6). Furthermore, we estimated the cleavage sites of TD1, which primarily cleaved the 19th and 21st phosphodiester bonds from the designated PAM position. These cleavage sites indicated that TD-1 cleaved the target double-stranded DNA with a three-base overhang at the 5' end (Figure 7).

[0060] To further investigate the sequence specificity of the PAM in the DNA used as the substrate plasmid for the cleavage reaction, we used one of the recognized PAMs, 5'-GATCA-3', and created sequences in which the first, second, third, and fourth bases from the 3' end were changed to four different bases. The cleavage efficiencies were then compared. As shown in Figure 8, while there were some differences in efficiency, cleavage was clearly observed at positions 1, 3, and 4. However, at position 2, cleavage efficiency was significantly reduced with bases other than C, resulting in almost no cleavage. This result also suggests that the PAM of TD1 may be determined solely by the C at position 2. Based on these results, we set the PAM position to 5'-CCTCA-3', and performed reactions using 270 nM TD1-crRNA and 5.7 nM target DNA in rCutsmart buffer (50 mM Potassium Acetate, 20 mM Tris-acetate, 10 mM Magnesium Acetate, 100 μg / ml Recombinant Albumin, pH 7.9 at 25°C) (New England Biolabs). After 20 minutes, much of the DNA remained uncleaved, but after 30 minutes, complete cleavage was observed. Therefore, we performed the same reaction conditions at different temperatures. The cleavage reaction was performed for 60 minutes. The results showed that the cleavage reaction proceeded efficiently between 37 and 45°C (Figure 9). Almost no cleavage occurred at 4°C or above 50°C, while many open circles (OCs), where only one strand was cleaved, were observed at 20 to 30°C (Figure 10).

[0061] Construction of TD1 Mutants To generate mutants lacking TD1 cleavage activity, we predicted the active site and performed site-directed mutagenesis to prepare the desired mutant protein. Specifically, we generated a mutant protein in which aspartic acid at position 308, glutamic acid at position 470, and aspartic acid at position 558, which are predicted to be the active site of the nuclease, were converted to alanine (Figure 10). This protein was mixed with the target DNA-crRNA shown in Figure 10 and subjected to gel filtration in the same manner as wild-type TD1. The TD1 mutant, crRNA, and target DNA were eluted in peak [1]. Cryo-electron microscopy of this fraction yielded particle images suitable for structural analysis. 1,248,508 particles were selected and grouped into nine groups, and 2D average images were obtained (Figure 11). Based on this data, a 3D model was constructed, and a model structure was obtained at 2.8 Å resolution (Figure 11). This structure contained crRNA and target DNA, and captured the state in which crRNA invaded the double-stranded target DNA to form a hybrid strand (R-loop formation) (Figure 12). Furthermore, we were able to identify the amino acids that recognize PAM. We also constructed and examined mutants, and confirmed that mutants in which arginine 191 and arginine 193 were replaced with alanine (R191A and R193A), respectively, exhibited almost no cleavage activity (Figure 13). All publications, patents, and patent applications cited herein are incorporated herein by reference in their entirety.

[0062] The present invention can be used for genome editing.

[0063] SEQ ID NO: 1: Sequence of the gene region encoding the novel CRISPR-Cas

Claims

1. A protein selected from the group consisting of the following (a), (b) and (c): (a) a protein consisting of the amino acid sequence of SEQ ID NO: 3; (b) a protein consisting of an amino acid sequence having an identity of 62% or more but less than 100% with the amino acid sequence of SEQ ID NO: 3, capable of forming a complex with crRNA and binding to DNA having a sequence complementary to the spacer sequence in crRNA; (c) a protein consisting of a sequence in which 1 to 20 amino acids are mutated in the amino acid sequence of SEQ ID NO: 3, capable of forming a complex with crRNA and binding to DNA having a sequence complementary to the spacer sequence in crRNA.

2. The protein according to claim 1, which has nuclease activity.

3. The protein according to claim 1, which has no nuclease activity.

4. A nucleic acid encoding the protein of claim 1.

5. An expression vector comprising DNA encoding the protein according to claim 1.

6. A cell transformed with the expression vector according to claim 5.

7. A method for producing any one of the following proteins (a), (b) or (c), comprising culturing the cell according to claim 6: (a) a protein consisting of the amino acid sequence of SEQ ID NO: 3; (b) a protein consisting of an amino acid sequence having an identity of 62% or more but less than 100% with the amino acid sequence of SEQ ID NO: 3, capable of forming a complex with crRNA and binding to DNA having a sequence complementary to the spacer sequence in crRNA; (c) a protein consisting of a sequence in which 1 to 20 amino acids have been mutated in the amino acid sequence of SEQ ID NO: 3, capable of forming a complex with crRNA and binding to DNA having a sequence complementary to the spacer sequence in crRNA.

8. A method for cleaving DNA having a sequence complementary to the spacer sequence in the crRNA using the protein and crRNA described in claim 1.

9. A method for genome editing using the protein and crRNA described in claim 1.

10. A kit for genome editing comprising at least one selected from the group consisting of the protein described in claim 1, mRNA encoding the protein described in claim 1, and an expression vector comprising DNA encoding the protein described in claim 1.

11. A crRNA having a repeat sequence comprising any one of the nucleotide sequences of SEQ ID NOs: 4 to 8.

12. The protein according to claim 1, which comprises an amino acid sequence homologous to the amino acid sequence of SEQ ID NO:3, and in which acidic or basic amino acids at positions corresponding to Asp156, Arg191, and Arg193 in the amino acid sequence of SEQ ID NO:3 form hydrogen bonds with bases in the PAM sequence of the target DNA.

13. The crRNA of claim 11, wherein the spacer forms an R-loop with the target DNA.

Citation Information

Patent Citations

  • Information processing system, method and program

    JP2023148437A

  • Compositions containing nucleases and uses thereof

    JP2023539569A

  • NOVEL Cas PROTEIN

    JP2024043578A

  • Novel crispr DNA targeting enzymes and systems

    WO2020018142A1

  • Novel crispr DNA targeting enzymes and systems

    WO2020180699A1