Novel CRISPR-associated proteins and uses thereof

Novel Cas12a proteins with amino acid substitutions improve genome editing precision and efficiency, addressing limitations of existing CRISPR-Cas systems by enhancing target recognition and cleavage, and offering potential cancer treatment applications.

JP7727042B2Active Publication Date: 2025-08-20G FLAS LIFE SCIENCES INC +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024064761
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-08-09
Filing Date
2024-04-12
Publication Date
2025-08-20
Estimated Expiration
2039-08-09

AI Technical Summary

Technical Problem

Existing CRISPR-Cas systems, such as Cas9 and Cas12a, face limitations in precision and efficiency for genome editing, particularly in recognizing target nucleic acid sequences and generating specific cleavage sites.

Method used

Development of novel Cas12a proteins with specific amino acid substitutions, such as lysine (Lys) at position 925 or 930, and aspartic acid (Asp) at positions 877 or 873, enhancing the protein's ability to recognize and cleave target nucleic acid sequences with improved precision and reduced non-specific activity.

Benefits of technology

The modified Cas12a proteins demonstrate enhanced endonuclease activity, enabling precise and diverse gene editing with reduced non-specific DNase activity, and can be used in pharmaceutical compositions for cancer treatment targeting specific nucleic acid sequences in cancer cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007727042000008
    Figure 0007727042000008
  • Figure 0007727042000009
    Figure 0007727042000009
  • Figure 0007727042000010
    Figure 0007727042000010
Patent Text Reader

Abstract

To provide a novel CRISPR-associated protein and a use thereof.SOLUTION: A protein represented by a specific amino acid sequence, according to the present invention, exhibits the activity of endonucleases, which recognize and cleave an intracellular nucleic acid sequence linked to a guide RNA. Therefore, a CRISPR-associated protein of the invention can be used as a different nuclease for genome editing, in a CRISPR-Cas system.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to novel CRISPR-associated proteins and uses thereof. [Background technology]

[0002] Genome editing is a technology that allows the free editing of genetic information in living organisms. Advances in the field of life sciences and the development of genome sequencing technology have made it possible to understand genetic information on a broad scale. For example, understanding the genes that govern reproduction, disease, and growth in plants and animals, genetic mutations that cause various human genetic diseases, and the production of biofuels has already been achieved; however, further technological advances must be made to directly utilize this understanding to improve organisms and treat human diseases.

[0003] Genome editing technology can be used to change the genetic information of animals, including humans, plants, and microorganisms, and therefore its scope of application can be dramatically expanded. Genetic scissors, which are molecular tools designed and manufactured to precisely cut desired genetic information, play a key role in genome editing technology. Similar to next-generation sequencing technology, which is taking the field of gene sequencing to the next level, the use of genetic scissors is increasing the speed and scope of genetic information utilization and is becoming a key technology to create new industrial fields.

[0004] Genetic scissors that have been developed so far can be divided into three generations based on their appearance order: the first generation of genetic scissors is zinc finger nuclease (ZFN); the second generation of genetic scissors is transcription activator-like effector nuclease (TALEN); and the recent research on clustered regularly interspaced palindromic short repeats (CRISPR) / CRISPR-associated protein 9 (Cas9) is the third generation of genetic scissors.

[0005] CRISPRs are genetic loci containing multiple short tandem repeats found in approximately 40% of sequenced bacterial genomes and 90% of sequenced archaeal genomes. When the Cas9 protein complexes with two RNAs called CRISPR RNA (crRNA) and transactivating crRNA (tracrRNA), it forms an active endonuclease, thereby shredding foreign genetic elements in invading phages or plasmids to protect host cells. The crRNA is transcribed from the CRISPR element in the host genome previously occupied by the foreign invader.

[0006] RNA-guided nucleases derived from the CRISPR-Cas system provide tools for genome editing. In particular, research into techniques that can edit the genomes of cells and organelles using single-stranded guide RNA (sgRNA) and Cas proteins has been actively pursued. Recently, the Cpf1 protein [obtained from the genera Prevotella and Francisella 1] has been reported as another nuclease protein in the CRISPR-Cas system (B. Zetsche et al., 2015), providing even more options for genome editing. Summary of the Invention [Problem to be solved by the invention]

[0007] As a result of continuing efforts to develop proteins that are more effective in genome editing than known nucleases, the inventors discovered a novel CRISPR-associated protein that recognizes and cleaves target nucleic acid sequences, and completed the present invention.

[0008] It is therefore an object of the present invention to provide novel CRISPR-associated proteins that recognize and cleave target nucleic acid sequences. [Means for solving the problem]

[0009] To achieve the above-mentioned objectives, the present invention provides a Cas12a protein having the amino acid sequence of SEQ ID NO: 1.

[0010] Additionally, the present invention provides a Cas12a protein having the amino acid sequence of SEQ ID NO: 1, in which the lysine (Lys) at position 925 has been substituted with another amino acid.

[0011] Additionally, the present invention provides a Cas12a protein having the amino acid sequence of SEQ ID NO:3.

[0012] Additionally, the present invention provides a Cas12a protein having the amino acid sequence of SEQ ID NO: 3, in which the lysine (Lys) at position 930 has been substituted with another amino acid.

[0013] Additionally, the present invention provides a Cas12a protein having the amino acid sequence of SEQ ID NO: 1, in which the aspartic acid (Asp) at position 877 has been substituted with another amino acid.

[0014] Additionally, the present invention provides a Cas12a protein having the amino acid sequence of SEQ ID NO: 3, in which the aspartic acid (Asp) at position 873 has been substituted with another amino acid.

[0015] Additionally, the present invention provides a pharmaceutical composition for treating cancer, comprising as active ingredients: mgCas12a; and crRNA targeting a nucleic acid sequence specifically present in cancer cells. [Effects of the Invention]

[0016] According to the present invention, the protein represented by the amino acid sequence of SEQ ID NO: 1 or SEQ ID NO: 3 has endonuclease activity that recognizes and cleaves intracellular nucleic acid sequences bound to guide RNA. Therefore, the novel CRISPR-associated protein of the present invention can be used as another nuclease for genome editing in the CRISPR-Cas system. [Brief explanation of the drawings]

[0017] [Figure 1] A schematic diagram of the process of discovering Cas12a from metagenomes. [Figure 2A] FIG. 1 shows a phylogenetic tree of discovered Cas12a genes. [Figure 2B] FIG. 1 shows the structures of novel Cas12a and AsCas12a. [Figure 3] FIG. 1 shows the amino acid sequences of existing Cas12a and mgCas12a of the present invention, and the sequences were aligned using the ESPript program. [Figure 4] FIG. 1 shows the amino acid sequences of existing Cas12a and mgCas12a of the present invention, and the sequences were aligned using the ESPript program. [Figure 5] FIG. 1 shows the amino acid sequences of existing Cas12a and mgCas12a of the present invention, and the sequences were aligned using the ESPript program. [Figure 6] FIG. 1 shows the amino acid sequences of existing Cas12a and mgCas12a of the present invention, and the sequences were aligned using the ESPript program. [Figure 7] FIG. 1 shows the amino acid sequences of existing Cas12a and mgCas12a of the present invention, and the sequences were aligned using the ESPript program. [Figure 8] FIG. 1 shows the amino acid sequences of existing Cas12a and mgCas12a of the present invention, and the sequences were aligned using the ESPript program. [Figure 9A] This is a table obtained by comparing and summarizing the sequence information of Cas12a and mgCas12a of the present invention. [Figure 9B] This is a table obtained by comparing and summarizing the sequence information of Cas12a and mgCas12a of the present invention. [Figure 10] 10 shows the results obtained by determining the pH-dependent activity of mgCas12a according to the present invention, while crRNA#1 in FIG. 10 has the nucleotide sequence of SEQ ID NO: 25, and crRNA#2 in FIG. 11 has the nucleotide sequence of SEQ ID NO: 26. [Figure 11] 10 shows the results obtained by determining the pH-dependent activity of mgCas12a according to the present invention, while crRNA#1 in FIG. 10 has the nucleotide sequence of SEQ ID NO: 25, and crRNA#2 in FIG. 11 has the nucleotide sequence of SEQ ID NO: 26. [Figure 12] 10 shows the results obtained by determining the pH-dependent activity of mgCas12a according to the present invention, while crRNA#1 in FIG. 10 has the nucleotide sequence of SEQ ID NO: 25, and crRNA#2 in FIG. 11 has the nucleotide sequence of SEQ ID NO: 26. [Figure 13] FIG. 1 shows the target nucleic acid sequence and the position where the crRNA binds. [Figure 14] Figure 1 shows the results obtained by identifying the gene editing efficiency achieved by each protein (Mock, mgCas12a-1 and mgCas12a-2) when crRNAs against the genes CCR5 and DNMT1, respectively, are used. [Figure 15] Figure 1 shows the results obtained by identifying the gene editing efficiency achieved by each protein (FnCpf1, mgCas12a-1 and mgCas12a-2) when two crRNAs are used for each gene FucT14-1 and FucT14-2. [Figure 16A] FIG. 1 shows the results obtained by identifying the DNA cleavage activity of FnCas12a, WT mgCas12a-1, or WT mgCas12a-2 proteins. [Figure 16B] FIG. 1 shows the results obtained by identifying the DNA cleavage activity of FnCas12a, WT mgCas12a-1, or WT mgCas12a-2 proteins. [Figure 17]This figure shows the results obtained by identifying the non-specific DNase function of existing Cas12a (AsCas12a, FnCas12a, or LbCas12a) and novel Cas12a (WT mgCas12a-1, d_mgCas12a-1, WT mgCas12a-2, or d_mgCas12a-2). [Figure 18A] FIG. 1 shows the results obtained by identifying whether FnCas12a, WT mgCas12a-1, or WT mgCas12a-2 proteins have nonspecific DNase function without crRNA. [Figure 18B] FIG. 1 shows the results obtained by identifying whether FnCas12a, WT mgCas12a-1, or WT mgCas12a-2 proteins have nonspecific DNase function without crRNA. [Figure 19] FIG. 1 shows results obtained by identifying whether mgCas12a can use the existing Cas12a 5′ handle to perform DNA cleavage. [Figure 20A] FIG. 1 shows the DNA cleavage activity of FnCas12a, mgCas12a-1, or mgCas12a-2 proteins in divalent ions. [Figure 20B] FIG. 1 shows the DNA cleavage activity of FnCas12a, mgCas12a-1, or mgCas12a-2 proteins in divalent ions. DETAILED DESCRIPTION OF THE INVENTION

[0018] BEST MODE FOR CARRYING OUT THE INVENTION In an embodiment of the present invention, a novel Cas12a protein obtained from a metagenome is provided.

[0019] As used herein, the term "Cas12a" refers to a CRISPR-associated protein, sometimes referred to as Cpf1. Additionally, Cpf1 is an effector protein found in type V CRISPR systems. Cas12a, a single effector protein, is similar to Cas9, an effector protein found in type II CRISPR systems, and combines with crRNA to cleave target genes. However, the two act differently. Cas12a protein interacts with a single-stranded crRNA. Therefore, unlike Cas9, Cas12a protein does not require the simultaneous use of crRNA and trans-activating crRNA (tracrRNA) or the creation of a single-stranded guide RNA (sgRNA) by synthetically combining tracrRNA and crRNA.

[0020] In addition, unlike Cas9, the Cas12a system recognizes the PAM located at the 5' position of the target sequence. In addition, the guide RNA that determines the target in the Cas12a system is also shorter than that in Cas9. In addition, Cas12a has the advantage that it generates a 5' overhang (sticky end) rather than a blunt end at the cleavage site in the target DNA, thereby enabling more precise and diverse gene editing.

[0021] Conventionally, Cas12a proteins can be obtained from the genera Candidatus, Lachnospira, Butyrivibrio, Peregrinibacteria, Acidaminococcus, Porphyromonas, Prevotella, Francisella, Candidatus Methanoplasma, or Eubacterium. In particular, PbCas12a is a protein obtained from Parcubacteria bacterium GWC2011_GWC2_44_17; PeCas12a is a protein obtained from Peregrinibacteria bacterium GW2011_GWA_33_10; AsCas12a is a protein obtained from Acidaminococcus sp. BVBLG; PmCas12a is a protein obtained from Porphyromonas macacae; LbCas12a is a protein obtained from Lachnospiraceae bacterium ND2006; PcCas12a is a protein obtained from Porphyromonas creviolicanis. PdCas12a is a protein obtained from Prevotella disiens; and FnCas12a is a protein obtained from Francisella novicida U112. However, each Cas12a protein may have different activities depending on the microorganism from which it is derived.

[0022] In the present invention, a novel Cas12a was identified by analyzing genes in a metagenome. Hereinafter, the metagenome-derived Cas12a may be referred to as mgCas12a. Similar to AsCas12a, the mgCas12a of the present invention contains the WED, REC, PI, RuvC, BH, and NUC domains (Figure 2). Additionally, similar to previously known Cas12a proteins, the mgCas12a protein of the present invention was identified as capable of performing gene cleavage using a crRNA and a gRNA containing a 5' handle. mgCas12a was identified to use a 5' handle RNA with the same sequence as FnCas12a. In particular, the 5' handle RNA may have the sequence AAUUUCUACUGUUGUAGAU (SEQ ID NO: 12). However, mgCas12a was also identified to function with the 5' handle RNAs in AsCas12a and LbCas12a (Figure 19).

[0023] mgCas12a may further comprise a tag for separation and purification. The tag may be attached to the N-terminus or C-terminus of mgCas12a. In addition, tags may be attached to the N-terminus and C-terminus of mgCas12a simultaneously. A specific example of a tag may be a 6xHis tag.

[0024] A specific example of mgCas12a is a protein having the amino acid sequence of SEQ ID NO: 1. Additionally, deletions or substitutions of amino acid portions may be made therein as long as the activity of mgCas12a is unchanged. Specifically, mgCas12a may be a protein having the amino acid sequence of SEQ ID NO: 1 in which lysine (Lys) at position 925 is substituted with another amino acid. Here, the other amino acid may be any one selected from the group consisting of arginine (Arg), histidine (His), aspartic acid (Asp), glutamic acid (Glu), serine (Ser), threonine (Thr), asparagine (Asn), glutamine (Gln), tyrosine (Tyr), alanine (Ala), isoleucine (Ile), leucine (Leu), valine (Val), phenylalanine (Phe), methionine (Met), tryptophan (Trp), glycine (Gly), proline (Pro), and cysteine (Cys). In particular, the protein may have the amino acid sequence of SEQ ID NO: 1 in which the lysine at position 925 has been substituted with glutamine, i.e. the protein may have the amino acid sequence of SEQ ID NO: 5.

[0025] Additionally, the gene encoding the protein having the amino acid sequence of SEQ ID NO: 1 may be a polynucleotide having the nucleotide sequence of SEQ ID NO: 2. Additionally, mgCas12a having the amino acid sequence of SEQ ID NO: 1 according to the present invention may have optimal activity at pH 7.0 to pH 7.9.

[0026] Another specific example of mgCpf1 is a protein having the amino acid sequence of SEQ ID NO: 3. Additionally, partial deletions or substitutions of amino acids may be made therein, as long as the activity of mgCpf1 is not altered. In particular, mgCpf1 may be a protein having the amino acid sequence of SEQ ID NO: 3 in which lysine (Lys) at position 930 is substituted with another amino acid. Here, the other amino acid may be any one selected from the group consisting of arginine (Arg), histidine (His), aspartic acid (Asp), glutamic acid (Glu), serine (Ser), threonine (Thr), asparagine (Asn), glutamine (Gln), tyrosine (Tyr), alanine (Ala), isoleucine (Ile), leucine (Leu), valine (Val), phenylalanine (Phe), methionine (Met), tryptophan (Trp), glycine (Gly), proline (Pro), and cysteine (Cys). In particular, the protein may have the amino acid sequence of SEQ ID NO: 3, in which the lysine at position 930 has been substituted with glutamine, i.e. the protein may have the amino acid sequence of SEQ ID NO: 6.

[0027] The gene encoding the protein having the amino acid sequence of SEQ ID NO:3 may be a polynucleotide having the nucleotide sequence of SEQ ID NO:4.

[0028] Additionally, mgCas12a having the amino acid sequence of SEQ ID NO: 3 according to the present invention may have optimal activity at pH 7.0 to pH 7.9.

[0029] In another aspect of the present invention, mgCas12a proteins with reduced endonuclease activity are provided. A specific example thereof may be mgCas12a having the amino acid sequence of SEQ ID NO: 1, in which aspartic acid (Asp) at position 877 has been substituted with another amino acid. The other amino acid may be any one selected from the group consisting of arginine (Arg), histidine (His), glutamic acid (Glu), serine (Ser), threonine (Thr), asparagine (Asn), glutamine (Gln), tyrosine (Tyr), alanine (Ala), lysine (Lys), isoleucine (Ile), leucine (Leu), valine (Val), phenylalanine (Phe), methionine (Met), tryptophan (Trp), glycine (Gly), proline (Pro), and cysteine (Cys). In particular, the protein may be a protein obtained by substituting aspartic acid (Asp) with alanine (Ala).

[0030] Another specific example of a mgCas12a protein is mgCas12a having the amino acid sequence of SEQ ID NO: 3, in which the aspartic acid (Asp) at position 873 has been replaced with another amino acid. The other amino acid may be any one selected from the group consisting of arginine (Arg), histidine (His), glutamic acid (Glu), serine (Ser), threonine (Thr), asparagine (Asn), glutamine (Gln), tyrosine (Tyr), alanine (Ala), lysine (Lys), isoleucine (Ile), leucine (Leu), valine (Val), phenylalanine (Phe), methionine (Met), tryptophan (Trp), glycine (Gly), proline (Pro), and cysteine (Cys). In particular, the protein may be a protein obtained by replacing aspartic acid (Asp) with alanine (Ala). Here, mgCas12a with reduced endonuclease activity may be referred to as inactivated mgCas12a or d_mgCas12a. d_mgCas12a may have the amino acid sequence of SEQ ID NO: 13 or SEQ ID NO: 14.

[0031] In addition, in yet another aspect of the present invention, a pharmaceutical composition for treating cancer is provided, comprising mgCas12a as an active ingredient and a crRNA targeting a nucleic acid sequence specifically present in cancer cells. Here, mgCas12a may have any one of the amino acid sequences selected from the group consisting of SEQ ID NO:1, SEQ ID NO:3, SEQ ID NO:5, and SEQ ID NO:6. As used herein, the term "nucleic acid sequence specifically present in cancer cells" refers to a nucleic acid sequence that is present only in cancer cells and not in normal cells. That is, this term refers to a sequence that differs from the sequence in normal cells, and the two sequences may differ by at least one nucleic acid. In addition, such a difference may be due to the substitution or deletion of a portion of a gene. In a specific example, the nucleic acid sequence specifically present in cancer cells may be a SNP present in cancer cells. A target DNA having the above-mentioned sequence present in cancer cells and a guide RNA having a sequence complementary to the target DNA specifically bind to each other.

[0032] In particular, with regard to nucleic acid sequences specifically present in cancer cells, crRNA can be created by discovering specific SNPs present only in cancer cells through genome sequencing of various cancer tissues and using those SNPs. This is done in a way that exhibits cancer cell-specific toxicity, thus enabling the development of patient-specific anti-cancer therapeutic drugs. In addition, nucleic acid sequences specifically present in cancer cells may be genes that have high copy number variations (CNVs) in cancer cells, unlike normal cells.

[0033] A specific example of cancer may be any one selected from the group consisting of bladder cancer, bone cancer, blood cancer, breast cancer, melanoma, thyroid cancer, parathyroid cancer, bone marrow cancer, rectal cancer, throat cancer, laryngeal cancer, lung cancer, esophageal cancer, pancreatic cancer, stomach cancer, tongue cancer, skin cancer, brain tumor, uterine cancer, head and neck cancer, gallbladder cancer, oral cancer, colon cancer, perianal cancer, central nervous system tumor, liver cancer, and colorectal cancer. In particular, the cancer may be stomach cancer, colorectal cancer, liver cancer, lung cancer, or breast cancer, which are known as the five major cancers in Korea.

[0034] Here, the crRNA targeting a nucleic acid sequence specifically present in cancer cells may include one or more gRNA sequences. For example, the crRNA may use a gRNA that can simultaneously target exons 10 and 11 of BRCA1 present in ovarian cancer or breast cancer. In addition, the crRNA may use two or more gRNAs targeting exon 11 of BRCA1. Therefore, the combination of gRNAs can be appropriately selected depending on the purpose of cancer treatment and the type of cancer. That is, different gRNAs may be selected and used. [Example]

[0035] [Mode of Invention] Hereinafter, the present invention will be described in more detail by the following examples. However, the following examples are for illustrative purposes only and the scope of the present invention is not limited thereto.

[0036] Example 1. Discovery of Metagenome-Derived Cas12a Protein Metagenome nucleotide sequences were downloaded from the NCBI Genbank BLAST database, and a local BLASTp database was constructed. In addition, the amino acid sequences of 16 Cas12a and various CRISPR-associated proteins (Cas1) were downloaded from the Uniprot database. The MetaCRT program was used to find CRISPR repeat and spacer sequences in the metagenome. Then, only metagenome sequences containing CRISPR sequences were extracted, and their genes were predicted using the Prodigal program.

[0037] Among the predicted genes, genes located 10 kb upstream or downstream of the CRISPR sequence were extracted, and Cas12a homologs were predicted from among the genes in question using the amino acid sequence of Cas12a. The Cas1 gene was used to predict whether a Cas1 homolog was located upstream or downstream of a Cas12a homolog; Cas12a genes between 800 and 1500 aa that were flanked by Cas1 were selected. For each of these genes, BLASTp was used in the NCBI Genbank non-redundant database to determine whether the gene was a previously reported gene or whether the gene was completely unrelated to CRISPR.

[0038] After removing fragmented Cas12a genes that did not start with methionine (Met), these genes were aligned using multiple alignment with the fast Fourier transform (MAFFT) program. Then, MEGA7 was used to construct a phylogenetic tree by neighbor-joining (100x bootstrap). Genes that formed a monophyletic group with previously known Cas12a genes were selected, and their evolutionary relationships were investigated by constructing a phylogenetic tree together with existing Cas12a amino acid sequences using MEGA7, maximum likelihood, and 1000x bootstrap. The process of discovering Cas12a from metagenomes is illustrated diagrammatically in Figure 1. Additionally, the Cas12a phylogenetic tree is illustrated in Figure 2A. Here, the novel protein with the amino acid sequence of SEQ ID NO: 1 was designated WT mgCas12a-1. Additionally, the novel protein with the amino acid sequence of SEQ ID NO: 3 was designated WT mgCas12a-2. In addition, the structures of AsCas12a, mgCas12a-1, and mgCas12a-2 are illustrated in Figure 2B.

[0039] Example 2. Creation of mgCas12a variants Cas12a candidates were aligned based on the structures of AsCas12a and LbCas12a using the ESPript program. For WT mgCas12a-1 and WT mgCas12a-2, partial amino acid substitutions were made to enhance their endonuclease activity. WT mgCas12a-1, in which the 925th amino acid, Lys (K), was replaced with Glu (Q), was named mgCas12a-1. Additionally, WT mgCas12a-2, in which the 930th amino acid, Lys (K), was replaced with Glu (Q), was named mgCas12a-2. The resulting variants were codon-optimized based on the codon usage of humans, Arabidopsis, and E. coli, and then gene synthesis was requested from Bionics. The nucleotide sequences of mgCas12a-1 and mgCas12a-2, optimized for human codons, are shown in SEQ ID NOs: 7 and 8, respectively. In addition, the amino acid sequences of existing Cas12a [AsCas12a (SEQ ID NO: 9), LbCas12a (SEQ ID NO: 10), and FnCas12a (SEQ ID NO: 11)] and Cas12a candidates (mgCas12a-1 and mgCas12a-2) aligned using the ESPript program are illustrated in Figures 3 to 8; the results obtained by comparing and summarizing the sequence information are illustrated in Figures 9A and 9B.

[0040] The WT mgCas12a-1, WT mgCas12a-2, mgCas12a-1, and mgCas12a-2 genes cloned in the pUC57 vector were then reinserted into the pET28a-KanR-6xHis-BPNLS vector, followed by cloning. The cloned vectors were transformed into E. coli strains DH5a and Rosetta, respectively. The 5' handle sequences of the crRNAs were extracted from the metagenomic CRISPR repeat sequences. The extracted RNA was synthesized into DNA oligos. The DNA oligos were transcribed using the MEGAshortscript T7 RNA Transcriptase Kit, and the concentration of the transcribed 5' handles was confirmed using a FLUOstar Omega.

[0041] Example 3. Protein expression and purification Five mL of overnight cultured E. coli Rosetta (DE3) was inoculated into 500 mL of liquid TB medium supplemented with 100 mg / mL kanamycin antibiotic. The medium was grown in an incubator at 37°C until the OD600 reached 0.6. For protein expression, the medium was treated with 0.4 μM isopropyl β-D-1-thiogalactopyranoside (IPTG), followed by further cultivation at 22°C for 16–18 hours. After centrifugation, the resulting cells were mixed with 10 mL of lysis buffer (20 mM HEPES pH 7.5, 100 mM KCl, 20 mM imidazole, 10% glycerol, and EDTA-free protease inhibitor cocktail) and then sonicated to disrupt the cells. The lysate was centrifuged three times for 20 minutes each at 6,000 rpm and then filtered through a 0.22 micron filter.

[0042] Subsequently, washing and elution were performed using a nickel column (HisTrap FF, 5 mL) and 300 mM imidazole buffer, and the protein was purified by affinity chromatography. Protein size was confirmed by SDS-PAGE electrophoresis, and dialysis was performed overnight against dialysis buffer (20 mM HEPES pH 7.5, 100 mM KCl, 1 mM DTT, 10% glycerol). The protein was then selectively subjected to filtration and concentration depending on its size (Amicon Ultra Centrifugal Filter 100,000 MWCO). In the case of proteins, their concentration was measured using the Bradford quantification method. The protein was then stored at -80°C until use.

[0043] Example 4. Identification of a suitable pH range for mgCas12a by cleavage analysis Lettuce (Lactuca sativa) xylosyltransferase was amplified by PCR to predict the protospacer adjacent motif (PAM), and guide RNAs (gRNAs) were designed based on the results. For mgCas12a-1 and mgCas12a-2 ribonucleoprotein (RNP) complexes, each mgCas12a protein was mixed with gRNA at a molecular ratio of 1:1.25 for 20 minutes at room temperature to prepare RNP complexes. Purified xylosyltransferase PCR products were then subjected to RNP treatment at various concentrations. Concentration adjustments were then made to NEBuffer 1.1 (1x buffer components, 10 mM Bis-Tris-propane-HCl, 10 mM MgCl, and 100 μg / mL BSA), NEBuffer 2.1 (1x buffer components, 50 mM NaCl, 10 mM Tris-HCl, 10 mM MgCl, and 100 μg / mL BSA), and NEBuffer 3.1 (1x buffer components, 100 mM NaCl, 50 mM Tris-HCl, 10 mM MgCl, and 100 μg / mL BSA), and in vitro cleavage assays were performed at 37°C, where NEBuffer 1.1, NEBuffer 2.1, and NEBuffer 3.1 had pH values of 7.0, 7.9, and 7.9, respectively, at 25°C. After each reaction was completed, the reaction was stopped by incubation at 65°C for 10 minutes, and the completed reaction was confirmed by 1.5% agarose gel electrophoresis. The results are illustrated in Figures 10 to 12. In Figures 10 to 12, mgCas12a-1 and mgCas12a-2 are designated hemgCas12a-1 and hemgCas12a-2, respectively. In addition, the target nucleic acid sequence in the xylosyltransferase and the position where the crRNA binds are indicated in the diagram, and this diagram is illustrated in Figure 13.

[0044] As shown in Figures 10 to 12, when the complex of mgCas12a-1 and crRNA was treated with NEBuffer 1.1, the target dsDNA was cleaved. In addition, when the complex of mgCas12a-2 and crRNA was treated with NEBuffer 1.1, the target dsDNA was cleaved. These results demonstrate that mgCas12a-1 and mgCas12a-2 are active at pH 7.0. Example 5. Analysis of gene editing efficiency of mgCas12a in animal cells Example 5.1. Generation of RNPs containing mgCas12a-1 or mgCas12a-2 for gene editing of CCR5 and DNMT1

[0045] HEK 293T cells were cultured in DMEM medium supplemented with 10% fetal bovine serum (FBS) and penicillin-streptomycin (P / S) at 37°C in a 5% CO2 incubator. RNPs were prepared by incubating 100 pmoles each of mgCas12a-1 and mgCas12a-2 proteins with 200 pmoles each of CCR5-targeting crRNA and DNMT1-targeting crRNA at room temperature for 20 minutes. The CCR5 and DNMT1 crRNA sequences were synthesized by Integrated DNA Technologies (IDT) and are shown in Table 1 below.

[0046] [Table 1]

[0047] Cultured HEK293T cells 2 x 10 5 Each cell was mixed with 20 μL of nucleofection reagent and then mixed with 10 μL of RNP complex. Transfection was then performed using a 4D-Nucleofector device (Lonza). At 48 and 72 hours post-transfection, genomic DNA was extracted from the cells using the PureLink™ Genomic DNA Mini Kit (Invitrogen).

[0048] Example 5.2. Sequencing Analysis of Target Sites The genomic DNA extracted in Example 5.1 was amplified using adapter primers for CCR5 or DNMT1 shown in Table 2 below.

[0049] [Table 2]

[0050] Purification and sequencing library preparation were then performed according to Illumina's protocol, and deep sequencing analysis was then performed on the target sites using a MiniSeq instrument. The gene editing efficiency achieved by mgCas12a-1 and mgCas12a-2 proteins is illustrated in Figure 14, and the sequencing analysis results of the target sites are shown in Table 3 below. As illustrated in Figure 14, mgCas12a-1 and mgCas12a-2 proteins exhibited higher gene editing efficiency than Mock protein.

[0051] [Table 3]

[0052] Example 6. Analysis of gene editing efficiency of mgCas12a in plant cells Example 6.1. Isolation of Plant Protoplasts Tobacco seeds were sterilized by treatment with 50% Clorox for 1 minute. The sterilized seeds were placed on a medium for seed germination and grown for 1 week. The seeds were then transferred to magenta boxes used for cultivation and grown for 3 weeks. The lighting conditions used were 16 hours of light and 8 hours of darkness, and the seeds were grown at temperatures of 25°C to 28°C. For plants, leaves grown for 4 to 6 weeks were used. The leaves were placed on a glass plate, and the leaf tips and stems were cut off so that only the inner parts of the leaves were used. Here, the leaves were cut into 0.5 mm or smaller pieces. The cut leaf pieces were placed in 10 mL of enzyme solution and incubated at room temperature in the dark on a rotary shaker (50 rpm) for 3 to 4 hours.

[0053] After incubation, 10 mL of W5 solution was added and mixed carefully. The protoplasts present in the enzyme solution were filtered using a cell strainer (70 μm). The filtered protoplasts were centrifuged at 100 × g for 6 minutes. The supernatant was discarded, and the protoplast pellet was carefully suspended by adding MMG solution. The suspension was then placed on ice for 10 to 30 minutes. The number of protoplasts in a portion of the suspension was counted using a Hem hemocytometer, a counting plate, and a microscope. Afterwards, when the protoplast concentration reached 2 × 10 6 The MMG solution was further added to dilute the solution to 100 cells / mL. The compositions of the enzyme solution, MMG solution, and PEG solution are shown in Table 4 below.

[0054] [Table 4]

[0055] Example 6.2. Sequencing analysis of target sites and identification of editing efficiency therefor crRNA, mgCas12a protein, and NEBuffer1.1 were added to a 2 mL tube to a final volume of 20 μL, and the reaction was allowed to proceed at room temperature for 10 minutes. 200 μL (5 × 10) of the protoplasts obtained in Example 6.1 were added. 5 The mixture (100 cells) and the reacted crRNA and mgCas12 protein (volume 20 μL) were added to a tube (2 mL), mixed well, and then incubated in a clean bench for 10 minutes. Then, 220 μL of PEG solution, the same volume as the incubation volume, was added to it and mixed carefully. The mixture was incubated at room temperature for 15 minutes. Next, 840 μL of W5 solution was added to it and mixed well. After centrifugation at 100 × g for 2 minutes, the supernatant was discarded. The culture was then carried out in W5 solution for 2 days. The cells were then harvested, and DNA was extracted from them.

[0056] The extracted DNA was used to perform PCR on the target region, and then the target gene editing efficiency was identified by next-generation sequencing (NGS). The results are shown in Table 5 below. As shown in Table 5, the gene editing efficiency achieved by mgCas12a-1 protein was 1.8-fold higher than that of FnCpf1.

[0057] [Table 5]

[0058] In addition, the gene editing efficiency achieved by using two crRNAs against the tobacco FucT14 gene was identified for each protein. The results are illustrated in Figure 15. As illustrated in Figure 15, the gene editing efficiency achieved by mgCas12a-1 protein was two-fold higher than that of FnCpf1. Here, the crRNA and primer sequences for the target genes NbFucT14_1 and NbFucT14_2 are shown in Tables 6 and 7 below.

[0059] [Table 6]

[0060] [Table 7]

[0061] Example 7. Comparison of gene editing efficiency between FnCas12a and mgCas12a To form ribonucleoprotein (RNP) complexes consisting of FnCas12a, WT mgCas12a-1, or WT mgCas12a-2 protein and crRNA, 6 pmol of FnCas12a, WT mgCas12a-1, or WT mgCas12a-2 protein and 7.5 pmol of crRNA were mixed with NEB1.1 buffer and 1x distilled water for 30 minutes at room temperature. To identify dsDNA cleavage activity using crRNA-dependent Cas12a (FnCas12a, WT mgCas12a-1, or WT mgCas12a-2), 0.3 pmol of target dsDNA (linear or circular) was added, and the reaction proceeded for 2 hours at 37°C. HsCCR5, HsDNMT1, and HsEMX1 were used as DNA fragments. Additionally, the linear DNAs (SEQ ID NOs: 27 to 29) used in the experiments were PCR-purified products, and the circular DNAs (SEQ ID NOs: 30 to 32) were purified plasmids. SDS and EDTA (gel loading dye, NEB) were added, and the mixture was then stored at -20°C for 10 minutes to stop the reaction. Each DNA was loaded onto a 1% agarose gel and then subjected to electrophoresis to confirm the DNA cleavage activity attributed to FnCas12a, WT mgCas12a-1, or WT mgCas12a-2. The results are illustrated in Figure 16A (linear DNA) and Figure 16B (circular DNA). In Figure 16A and Figure 16B, S indicates substrate, and the numbers below the gel indicate the intensity of the substrate DNA band.

[0062] Example 8. Identification of non-specific DNase activity of mgCas12a To identify the random DNase function of Cas12a (AsCas12a, FnCas12a, or LbCas12a) and mgCas12a (WT mgCas12a-1, d_mgCas12a-1, WT mgCas12a-2, or d_mgCas12a-2), experiments were performed in the same manner as in Example 7. Here, d-mgCas12a-1 and d_mgCas12a-2 refer to proteins derived from WT mgCas12a-1 and WT mgCas12a-2, respectively, by substituting Asp (position 877 of WT mgCas12a-1 or position 873 of WT mgCas12a-2) with Ala.

[0063] Specifically, to form ribonucleoprotein (RNP) complexes consisting of each of the seven forms of Cas12a and crRNA, 6 pmol of each Cas12a protein and 7.5 pmol of crRNA were reacted at room temperature for 30 minutes in the presence of NEB1.1 buffer and 1x distilled water. Then, 0.3 pmol of target dsDNA was added, and the reaction proceeded at 37°C for 12 or 24 hours. Here, HsCCR5, HsDNMT1, and HsEMX1 were used as DNA. SDS and EDTA (gel loading dye, NEB) were added, and the mixture was then stored at -20°C for 10 minutes to terminate the reaction. Each DNA was loaded onto a 1% agarose gel and then subjected to electrophoresis to confirm the DNA cleavage activity attributed to the seven forms of Cas12a. The results are illustrated in Figure 17. In Figure 17, S represents substrate, and the numbers below the gel indicate the intensity of the substrate DNA band.

[0064] As shown in Figure 17, the ribonucleoprotein complexes consisting of the novel Cas12a (WT mgCas12a-1, d_mgCas12a-1, WT mgCas12a-2, or d_mgCas12a-2) and crRNA exhibited weaker nonspecific DNase activity than the ribonucleoprotein complexes consisting of the existing Cas12a (AsCas12a, FnCas12a, or LbCas12a) and crRNA. In addition, the overall reaction of Cas12a RNP with DNA may be expected to result in nonspecific DNase activity.

[0065] Example 9. Identification of the non-specific DNase function of Cas12a in the absence of crRNA To determine whether Cas12a possesses random DNase function for FnCas12a, WT mgCas12a-1, or WT mgCas12a-2 proteins even in the absence of crRNA, experiments were performed for varying times in the same manner as in Example 7, except that crRNA-free conditions were used. The results are illustrated in Figures 18A and 18B. As illustrated in Figures 18A and 18B, FnCas12a, WT mgCas12a-1, or WT mgCas12a-2 proteins possess random DNase function even in the absence of crRNA, and the random DNase function of FnCas12a protein was the first to emerge.

[0066] Example 10. Identification of the DNA cleavage function of mgCas12a using existing Cas12a handles To determine whether the new Cas12a (d_mgCas12a or WT mgCas12a) could perform DNA cleavage using the handle located at the 5' end of an existing Cas12a (AsCas12a, FnCas12a, or LbCas12a) sequence, experiments were performed in the same manner as in Example 7, except that the handles of AsCas12a, FnCas12a, or LbCas12a were used, with varying reaction times. The results are illustrated in Figure 19.

[0067] As shown in Figure 19, when DNA cleavage was performed using d_mgCas12a or WT mgCas12a proteins with AsCas12a, FnCas12a, or LbCas12a handles, the DNA cleavage efficiency varied slightly depending on the handle, but all d_mgCas12a or WT mgCas12a proteins using the three handles exhibited DNA cleavage function. These results demonstrate that mgCas12a can use AsCas12a, FnCas12a, or LbCas12a handles for DNA cleavage.

[0068] Example 11. Identification of the activity of FnCas12a or mgCas12a in divalent ions Additionally, to identify the DNA cleavage activity of FnCas12a, mgCas12a-1, or mgCas12a-2 proteins in divalent ions (CaCl, CoCl, CuSO, FeCl, MnSO, NiSO, or ZnSO), experiments were performed in the same manner as in Example 4, except that predetermined amounts of divalent ions were used instead of NEBuffer 1.1. The results are illustrated in Figures 20A and 20B. As illustrated in Figures 20A and 20B, FnCas12a, mgCas12a-1, or mgCas12a-2 proteins exhibited similar DNA cleavage activity in the same divalent ions.

Claims

[Claim 1] Cas12a protein having the amino acid sequence of SEQ ID NO:1.

Citation Information

Patent Citations

  • Compositions and methods for modifying genomes

    US20170233756A1

  • Novel engineered and chimeric nucleases

    WO2018071672A1