DNA cleavage means based on the Cas9 protein derived from the bacterium Clostridium cellulolyticum, which is important in biotechnology

The novel CRISPR-Cas9 system from Clostridium cellulolyticum with a two-letter PAM sequence and broad temperature range addresses limitations in existing CRISPR-Cas systems, enabling precise and efficient genomic editing across diverse organisms.

JP7698578B2Active Publication Date: 2025-06-25JOINT CO BIOCAD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2021529802
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-11-26
Filing Date
2019-11-26
Publication Date
2025-06-25
Estimated Expiration
2039-11-26

AI Technical Summary

Technical Problem

Current CRISPR-Cas systems are limited by specific PAM sequences, restricting the use of double-strand breaks in DNA regions, and there is a need for tools that can modify genomic DNA sequences in diverse organisms with precision and efficiency.

Method used

The use of a novel CRISPR-Cas9 system from Clostridium cellulolyticum (CcCas9) with a two-letter PAM sequence (NNNNGNA) and a broad temperature range (37°C to 65°C) allows for precise double-strand breaks in DNA, facilitated by a guide RNA and a Cas9 protein with a unique amino acid sequence (SEQ ID NO: 1 or variants with 95% identity).

Benefits of technology

This system expands the versatility of CRISPR-Cas9 by enabling double-strand breaks at a greater number of specific sites and over a wider temperature range, enhancing genomic editing efficiency and versatility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698578000010
    Figure 0007698578000010
  • Figure 0007698578000011
    Figure 0007698578000011
  • Figure 0007698578000012
    Figure 0007698578000012
Patent Text Reader

Abstract

The present invention describes a novel bacterial nuclease from the CRISPR-Cas9 system derived from the bacterium Clostridium cellulolyticum and its use to create precisely specific double-stranded breaks in DNA molecules. This nuclease has unique properties and can be used as a tool to introduce modifications at precisely defined sites in the genomic DNA sequences of unicellular or multicellular organisms. This increases the versatility of available CRISPR-Cas9 systems, enabling the use of Cas9 nucleases from a variety of organisms to cut genomic or plasmid DNA at a greater number of specific sites and over a wider temperature range. Furthermore, it provides easier genome editing for the biotechnologically important bacterium Clostridium cellulolyticum.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of molecular biology and microbiology. In particular, the present invention discloses novel bacterial nucleases of the CRISPR-Cas system. The present invention may be used as a tool for the precise and specific modification of DNA in various organisms.

Background Art

[0002] The modification of DNA sequences is one of the current issues in the field of biotechnology today. Editing and modifying the genomes of eukaryotic and prokaryotic organisms, as well as the manipulation of DNA in vitro, require the targeted introduction of double-strand breaks into the DNA sequence. To solve this problem, currently, the following techniques are used: artificial nuclease systems containing zinc finger-type domains, TALEN systems, and the bacterial CRISPR-Cas system. The first two techniques require the effort of optimizing the nuclease amino acid sequence for the recognition of specific DNA sequences. In contrast, in the case of the CRISPR-Cas system, the structure that recognizes the DNA target is not a protein but a small molecule guide RNA. For the cleavage of a specific DNA target, there is no need to newly synthesize the nuclease or its gene, but it is performed by using a guide RNA complementary to the target sequence. For this reason, the CRISPR Cas system has become a suitable and efficient means for cleaving various DNA sequences. This technique enables the simultaneous cleavage of DNA in several regions using guide RNAs of different sequences. Such an approach is also used for simultaneously modifying several genes in eukaryotes.

[0003] Due to its properties, the CRISPR-Cas system is a prokaryotic immune system capable of introducing cleavage very specifically within viral genetic material (Mojica F. J. M. et al. Intervening sequences of regularly spaced prokaryotic repeats derive from foreign genetic elements / / Journal of molecular evolution. 2005. Vol. 60. Issue 2. pp.174-182). The abbreviation CRISPR-Cas stands for "Clustered Regularly Interspaced Short Palindromic Repeats and CRISPR associated Genes" (Jansen R. et al. Identification of genes that are associated with DNA repeats in prokaryotes / / Molecular microbiology. 2002. Vol. 43. Issue 6. pp.1565-1575). All CRISPR-Cas systems consist of a CRISPR cassette and genes encoding various Cas proteins (Jansen R. et al., Molecular microbiology. 2002. Vol. 43. Issue 6. pp.1565-1575). The CRISPR cassette consists of spacers, each having a unique nucleotide sequence, and repetitive palindromic repeats (Jansen R. et al., Molecular microbiology. 2002. Vol. 43. Issue 6. pp.1565-1575).Transcription of the CRISPR cassette, followed by its processing, results in the formation of guide crRNA, which, together with the Cas protein, forms an effector complex (Brouns S. J. J. et al. Small CRISPR RNAs guide antiviral defense in prokaryotes / / Science. 2008. Vol. 321. Issue 5891. pp.960-964). Through complementary pairing between the crRNA and the target DNA site called the protospacer, the Cas nuclease recognizes the DNA target and introduces a cleavage therein in a very specific manner.

[0004] CRISPR-Cas systems containing a single effector protein are grouped into six different types (types I-VI) depending on the Cas protein included in the system. The type II CRISPR-Cas9 system is characterized by a simple composition and mechanism of action, i.e., its function requires only the formation of an effector complex consisting of one Cas9 protein and two small RNAs as follows: crRNA and tracer RNA (tracrRNA). The tracer RNA forms complementary pairs with the crRNA region generated from the CRISPR repeat to form the secondary structure necessary for the guide RNA to bind to the Cas effector. The Cas9 effector protein has two nuclease domains (HNH and RuvC) that introduce cleavage within the complementary strand of the target DNA and is thus an RNA-dependent DNA endonuclease that forms a double-stranded DNA break (Deltcheva E. et al. CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III / / Nature. 2011. Vol. 471. Issue 7340. p.602).

[0005] To date, several CRISPR-Cas nucleases are known that are capable of targeting and specifically introducing double-strand breaks into DNA. One of the main characteristics that limits the use of the CRISPR-Cas system is the PAM sequence that is adjacent to the 3' end of the DNA target and whose presence is required for the accurate recognition of DNA by the Cas9 nuclease. Different CRISPR-Cas proteins have different PAM sequences, thus limiting the potential for the use of nucleases in any DNA region. To enable the modification of any DNA region, both in vitro and in the genomes of organisms, the use of CRISPR-Cas proteins with novel diverse PAM sequences is required. Modification of eukaryotic genomes also requires the use of small-sized nucleases to provide AAV-mediated delivery of the CRISPR-Cas system into cells.

[0006] Although many techniques are known for cutting DNA and modifying genomic DNA sequences, there remains a need for new and effective means for modifying DNA in diverse organisms and at precisely defined sites within the DNA sequence. The present invention provides several characteristics necessary to solve this problem.

[0007] The basis of the present invention is the CRISPR Cas system discovered in the bacterium Clostridium cellulolyticum. The anaerobic bacterium Clostridium cellulolyticum (C. cellulolyticum) is capable of hydrolyzing lignocellulose without the addition of commercial cellulases and forming lactate, acetate, and ethanol as end products (Desvaux M. Clostridium cellulolyticum: model organism of mesophilic cellulolytic clostridia. FEMS Microbiol Rev. 2005 Sep;29(4):741-64). Due to these capabilities of these microorganisms, they are promising candidates for biofuel producers. The use of production bacteria such as C. cellulolyticum in biotechnological production would assist in making the raw material processing cycle more efficient, increasing efficiency, and ultimately reducing the burden on all components of the biosphere. Genetic engineering methods can also be used to significantly improve microbial metabolic parameters and shift the balance to support the production of more butanol than lactate and acetate. For example, double mutants of the lactate and malate dehydrogenase genes showed the absence of lactate formation and an increase in butanol yield (Li Y et al., Combined inactivation of the Clostridium cellulolyticum lactate and malate dehydrogenase genes substantially increases ethanol yield from cellulose and switchgrass fermentations. Biotechnol Biofuels. 2012 Jan 4;5(1):2). So far, it has not been possible to develop an effective method for producing C. cellulolyticum strains with mutations in the phosphotransacetylase and acetate kinase genes that can reduce acetate production. The genome of Clostridium cellulolyticum as well as other organisms may be modified using the present invention.

Prior Art Documents

Non-Patent Documents

[0008]

Non-Patent Document 1

Non-Patent Document 2

Non-Patent Document 3

Non-Patent Document 4

Non-Patent Document 5

Non-Patent Document 6

Summary of the Invention

Problems to be Solved by the Invention

[0009] An object of the present invention is to provide a novel means for modifying genomic DNA sequences of single-celled or multi-celled organisms using the CRISPR-Cas9 system. Current systems have limitations in use due to specific PAM sequences that must be present at the 3' end of the DNA region to be modified. Searching for novel Cas9 enzymes with other PAM sequences will expand the range of available means for forming double-strand breaks at desired precise specific sites in DNA molecules of various organisms.

Means for Solving the Problems

[0010] To solve this problem, the authors characterized the type II CRISPR nuclease CcCas9, previously predicted for C. cellulolyticum, which can be used to introduce directed modifications within the genomes of both the above and other organisms. The essential features that characterize the present invention are as follows: (a) a small molecule two-letter PAM sequence different from other known PAM sequences; (b) a characteristic small size of the CcCas9 protein, which is 1030 amino acid residues (a.a.r.) and 23 a.a.r. smaller than that of the known Cas9 enzyme (SaCas9) from Staphylococcus aureus; (c) a broad operating temperature range of the CcCas9 nuclease, which is active at temperatures from 37°C to 65°C and optimal at 45°C, and will be usable in organisms with diverse temperatures.

[0011] The said problem is solved by using a protein comprising the amino acid sequence of SEQ ID NO: 1 or an amino acid sequence that is at least 95% identical to the amino acid sequence of SEQ ID NO: 1 and differs only at non-conserved amino acid residues from SEQ ID NO: 1, to form a double-strand break in the DNA molecule immediately preceding the nucleotide sequence 5'-NNNNGNA-3' in the DNA molecule. N is intended to refer to any nucleotide (A, G, C, T). In some embodiments of the present invention, this use is characterized by the formation of a double-strand break in the DNA molecule at a temperature from 37°C to 65°C. In a preferred embodiment of the present invention, this use is characterized by the formation of a double-strand break in the DNA molecule at a temperature from 37°C to 55°C.

[0012] The above problem is solved by using a method for modifying the genomic DNA sequence of a unicellular or multicellular organism, which comprises introducing into at least one cell of the organism an effective amount of: a) either a protein comprising the amino acid sequence of SEQ ID NO: 1 or a nucleic acid encoding a protein comprising the amino acid sequence of SEQ ID NO: 1, and b) either a guide RNA comprising a nucleotide sequence that forms a double strand with the nucleotide sequence of the genomic DNA region of the organism immediately adjacent to the nucleotide sequence 5'-NNNNGNA-3' and interacts with the protein after double strand formation, or a DNA sequence encoding the guide RNA, wherein the interaction of the protein with the guide RNA and the nucleotide sequence 5'-NNNNGNA-3' results in the formation of a double strand break in the genomic DNA sequence immediately adjacent to the sequence 5'-NNNNGNA-3'.

[0013] A mixture of crRNA and tracer RNA (tracrRNA) that can form a complex with the target DNA region and the CcCas9 protein may be used as the guide RNA. In a preferred embodiment of the invention, a hybrid RNA constructed based on crRNA and tracer RNA may be used as the guide RNA. Methods for constructing hybrid guide RNAs are known to those skilled in the art (Hsu PD et al., DNA targeting specificity of RNA-guided Cas9 nucleases. Nat Biotechnol. 2013 Sep;31(9):827-32).

[0014] The present invention may be used for both in vitro cleavage of target DNA and modification of the genomes of several organisms. The genome may be modified directly, i.e., by cleaving the genome at the corresponding site, and also by inserting an exogenous DNA sequence through homologous repair.

[0015] Any region of double-stranded or single-stranded DNA derived from a biological genome other than that used for administration (or a composition of such regions in themselves and those containing other DNA fragments) may be used as an exogenous DNA sequence, where the region (or composition of regions) is intended to be induced by CcCas9 nuclease and incorporated into the site of double-strand break in the target DNA. In some embodiments of the present invention, a double-stranded DNA region derived from a biological genome, which is used for the introduction of CcCas9 protein but is further modified by mutation (nucleotide substitution) and by insertion or deletion of one or more nucleotides, may be used as an exogenous DNA sequence.

[0016] The technical result of the present invention is to increase the versatility of the available CRISPR-Cas9 system and enable the use of Cas9 nuclease to cut genomic or plasmid DNA at a greater number of specific sites and over a wider temperature range.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Mode for Carrying Out the Invention

[0018] As used in the description of the present invention, the terms "includes" and "including" should be construed to mean "including, among other things". These terms are not intended to be construed as "consisting only of". Unless otherwise defined, technical and scientific terms in this application have the typical meanings generally recognized in scientific and technical literature.

[0019] As used herein, the term "percent homology of two sequences" is equivalent to the term "percent identity of two sequences". Sequence identity is determined based on a reference sequence. Algorithms for sequence analysis are known in the art, such as BLAST described in Altschul et al., J. Mol. Biol., 215, pp. 403-10 (1990). For the purposes of the present invention, to determine the level of identity and similarity between nucleotide sequences and amino acid sequences, comparison of nucleotide and amino acid sequences may be used, which is performed with standard parameters and using gapped alignment by the BLAST software package provided by the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / blast). The percent identity of two sequences is determined by the number of identical amino acid positions in these two sequences, taking into account the number of gaps inserted for the optimal comparison of the two sequences by alignment and the length of each gap. The percent identity is equal to the number of amino acids that are identical at a given position, divided by the total number of positions, considering the sequence alignment, and multiplied by 100.

[0020] The term "specifically hybridizes" refers to the association between two single-stranded nucleic acid molecules or sufficiently complementary sequences, which enables such hybridization under predetermined conditions typically used in the art.

[0021] The phrase "double-stranded break located immediately before in the nucleotide PAM sequence" means that the double-stranded break in the target DNA sequence is created at a distance of 0 to 25 nucleotides before the nucleotide PAM sequence.

[0022] A protein containing a specific amino acid sequence is intended to refer to a protein having the amino acid sequence and an amino acid sequence composed of other sequences that may be linked to the amino acid sequence by peptide bonds. Examples of other sequences may be a nuclear localization signal (NLS) or other sequences that provide an increase in the functionality of the amino acid sequence.

[0023] The exogenous DNA sequence introduced simultaneously with the guide RNA is intended to refer to a DNA sequence specifically prepared for the specific modification of double-stranded target DNA at the cleavage site determined by the specificity of the guide RNA. Such modifications may be, for example, insertions or deletions of specific nucleotides at the cleavage site in the target DNA. The exogenous DNA may be either a DNA region from a different organism or a DNA region from the same organism as that of the target DNA.

[0024] The effective amounts of the protein and RNA introduced into the cell are intended to refer to the amounts of the protein and RNA such that when introduced into the cell, they can form a functional complex, that is, a complex that specifically binds to the target DNA and causes a double-stranded break at the site determined by the guide RNA and the PAM sequence on the DNA. The efficiency of this process can be evaluated by analyzing the target DNA isolated from the cell using conventional techniques known to those skilled in the art.

[0025] Proteins and RNAs may be delivered into cells by a variety of techniques. For example, a protein may be delivered as a DNA plasmid encoding the gene of this protein, as mRNA translated into this protein in the cytoplasm, or as a ribonucleoprotein complex containing this protein and guide RNA. Delivery may be performed by a variety of techniques known to those skilled in the art.

[0026] The nucleic acid encoding the components of the system may be introduced directly or indirectly into cells by transfection or transformation of cells by methods known to those skilled in the art, by use of recombinant viruses, or by manipulation of cells, such as DNA microinjection, etc.

[0027] A ribonucleic complex consisting of a nuclease, guide RNA, and exogenous DNA (if necessary) may be delivered by transfecting the complex into cells or by mechanically introducing the complex into cells, such as by microinjection.

[0028] The nucleic acid molecule encoding the protein to be introduced into cells may be integrated into the chromosome or may be DNA replicated extrachromosomally. In some embodiments, in order to ensure efficient expression of the protein gene in the introduced DNA, since the occurrence frequencies of synonymous codons are uneven in the coding regions of various biological genomes, it is necessary to modify the sequence of the DNA according to the cell type so as to optimize the codons for expression. Codon optimization is necessary to increase expression in animal, plant, fungal, or microbial cells.

[0029] For a protein having a sequence that is at least 95% identical to the amino acid sequence of SEQ ID NO: 1 to function in a eukaryotic cell, it is necessary for this protein to ultimately reach the nucleus of this cell. Thus, in some embodiments of the present invention, a protein having a sequence that is at least 95% identical to the amino acid sequence of SEQ ID NO: 1 and further modified at one or both ends by the addition of one or more nuclear localization signals is used to form a double-strand break in the target DNA. For example, a nuclear localization signal derived from the SV40 virus may be used. To provide efficient delivery to the nucleus, for example, the nuclear localization signal may be separated from the major protein sequence by a spacer sequence as described in Shen B et al., "Generation of gene-modified mice via Cas9 / RNA-mediated gene targeting", Cell Res. 2013 May;23(5):720-3. Further, in other embodiments, different nuclear localization signals, or alternative methods for delivering the protein into the cell nucleus may be used.

[0030] The present invention includes the use of a protein from a Clostridium cellulolyticum (C. cellulolyticum) organism that is homologous to a previously characterized Cas9 protein for introducing a double-strand break within a DNA molecule at a precisely defined position.

[0031] In metabolic engineering, editing the C. cellulolyticum genome is a difficult task due to the lack of efficient editing tools. Methods for targeted genome editing, such as recombineering, group II intron retrotransposition, and allelic exchange, have several significant limitations. For example, the method of recombination-dependent allelic exchange is quite time-consuming and has low efficiency (Heap J. T. et al. Integration of DNA into bacterial chromosomes from plasmids without a counter-selection marker / / Nucleic acids research. 2012. Vol. 40. Issue 8. pp.e59-e59). Insertion of long DNA fragments, such as metabolic pathway transfer, is difficult with current genome editing tools that require the presence of recombination sites and / or recombinases (Esvelt K. M., Wang H. H. Genome-scale engineering for systems and synthetic biology / / Molecular systems biology. 2013. Vol. 9. Issue 1. p.641). A simple and efficient method is needed to successfully perform genome engineering and generate mutants with predefined characteristics.

[0032] The use of CRISPR nucleases to introduce targeted modifications into the genome has several advantages. First, the specificity of the system's activity is determined by the crRNA sequence, which allows the use of one type of nuclease for all target loci. Second, the technology enables the simultaneous delivery into the cell of several guide RNAs complementary to different gene targets, thereby allowing several genes to be modified simultaneously.

[0033] Furthermore, the use of the native CRISPR-Cas9 system from the bacterium C. cellulolyticum makes the system for editing the genome of this organism easier and more efficient, because this method does not require the introduction, maintenance, and expression of foreign genes into cells. Instead, it is possible to develop a method for introducing a guide RNA directed against the target gene into the bacterium, whereby the CRISPR-Cas9 system in the host cell would be able to recognize the corresponding DNA target of the bacterium important for biotechnology and introduce a double-strand break therein.

[0034] For the biochemical characterization of the Cas9 protein from C. cellulolyticum H10, the CRISPR locus encoding the main system components (CcCas9, cas1, cas2 protein genes, as well as the CRISPR cassette and guide RNA) was cloned into the single-copy bacterial vector pACYC184. The effector ribonucleic acid complex consisting of Cas9 and the crRNA / tracrRNA (trans-activating crRNA) duplex requires the presence of a PAM (protospacer adjacent motif) on the DNA target for DNA recognition and subsequent hydrolysis, in addition to crRNA spacer-protospacer complementarity (Mojica F. J. M. et al. Short motif sequences determine the targets of the prokaryotic CRISPR defence system / / Microbiology. 2009. Vol. 155. Issue 3. pp.733-740). The PAM is a strictly defined sequence of several nucleotides located within type II systems, adjacent to or a few nucleotides away from the 3’ end of the protospacer on the non-target strand. In the absence of the PAM, hydrolysis of DNA binding with double-strand break formation does not occur. The requirement for the presence of the PAM sequence on the target increases the recognition specificity but at the same time imposes a constraint on the selection of the target DNA region for introducing a break.

[0035] To determine the sequence of the guide RNA of the CRISPR-Cas9 system, RNA sequencing was performed on Escherichia coli (E. coli) DH5 alpha bacteria carrying the generated pACYC184_CcCas9 construct. Sequencing showed that the CRISPR cassette of the system was actively transcribed, as was the tracer RNA (Figure 1). Analysis of the crRNA and tracrRNA sequences made it possible to consider that they could form a secondary structure that is probably recognized by the CcCas9 nuclease.

[0036] Furthermore, the authors determined the PAM sequence of the CcCas9 protein using bacterial PAM screening. To determine the PAM sequence of the CcCas9 protein, E. coli DH5 alpha cells carrying the pACYC184_CcCas9 plasmid were transformed with a plasmid library containing the spacer sequence 5'-TAAAAAATAAGCAAGCGATGATATGAATGC-3' of the CRISPR cassette of the CcCas9 system, with a random 7-character sequence adjacent to the 5' or 3' end. Plasmids carrying sequences corresponding to the PAM sequence of the CcCas9 system were subjected to degradation under the action of a functional CRISPR-Cas system, while the remaining library plasmids were effectively transformed intracellularly and made resistant to the antibiotic ampicillin. After transformation and incubation of the cells on plates containing the antibiotic, the colonies were washed off the agar surface and DNA was extracted from them using the Qiagen Plasmid Purification Midi Kit. From the plasmid isolation pool, the region containing the randomized PAM sequence was amplified by PCR and then subjected to high-throughput sequencing on the Illumina platform. The resulting reads were analyzed by comparing the transformation efficiency of plasmids containing unique PAMs included in the library into cells carrying pACYC184_CcCas9 or control cells carrying the empty vector pacyc184. The results were analyzed using bioinformatics methods. As a result, it was possible to identify the PAM of the CcCas9 system, which is the two-letter sequence NNNNGNA (Figure 2).

[0037] Next, the PAM sequence was further determined by reproducing the cleavage reaction in vitro. To determine the PAM sequence of the CcCas9 protein, in vitro cleavage of a double-stranded PAM library was used. For this purpose, all components of the CcCas9 effector complex as follows: guide RNA and recombinant nuclease had to be obtained. Determination of the guide RNA sequence by RNA sequencing made it possible to synthesize crRNA and tracrRNA molecules in vitro. Synthesis was performed using the NEB HiScribe T7 RNA Synthesis Kit. The double-stranded DNA library was a 374 bp fragment containing a protospacer sequence flanked by randomized 7 nucleotides (5’NNNNNNN3’) at the 3’ end:

[0038]

Chemical formula

[0039] To cut this target, guide RNAs with the following sequences were used:

[0040]

Chemical formula

[0041] The bold type indicates the crRNA sequence complementary to the protospacer (target DNA sequence). To obtain the recombinant CcCas9 protein, its gene was cloned into the plasmid pET21a. The resulting plasmid CcCas9_pET21a was used to transform Escherichia coli Rosetta cells. Cells carrying the plasmid were grown to an optical density of OD_600 = 0.6, and then the expression of the CcCas9 gene was induced by adding IPTG to a concentration of 1 mM. The cells were incubated at 25 °C for 4 hours and then lysed. The recombinant protein was purified in two steps as follows: by affinity chromatography (NiNTA) and by protein size exclusion on a Superdex 200 column. The resulting protein was concentrated using an Amicon 30 kDa filter. Subsequently, the protein was frozen at -80 °C and used for in vitro reactions.

[0042] The in vitro reaction to cut the linear PAM library was carried out under the following conditions: 1xCutSmart buffer 400 nM CcCas9 100 nM DNA library 2 μM crRNA 2 μM tracrRNA The total reaction volume was 20 μl.

[0043] Clostridium cellulolyticum H10 is commonly found in compost piles and has an optimal cleavage temperature of 45 °C, and thus the reaction at this temperature was carried out for 30 minutes. As a result of the cut, the library fragment portion was split into two parts with lengths of approximately 50 base pairs (bp) and 324 bp. As a control sample, a reaction without the addition of crRNA, which is an essential component of the Cas effector complex, was used.

[0044] The reaction products were applied onto a 1.5% agarose gel and subjected to electrophoresis. The uncut DNA fragment of 374 bp in length was extracted from the gel and prepared for high-throughput sequencing using the NEB Next Ultra II kit. The samples were sequenced on an Illumina platform and then the sequences were analyzed using bioinformatics methods, where the differences in the nucleotide presentation at individual positions of the PAM (NNNNNNN) were determined as compared to the control samples (Figure 3).

[0045] As a result, the authors were able to determine the PAM sequence of CcCas9: NNNNGNA by an in vitro method, which completely replicated the results obtained from experiments using bacteria.

[0046] Next, the significance of individual positions of the PAM sequence was checked. For this purpose, an in vitro reaction was carried out when cutting a DNA fragment containing the DNA target 5’-gtgctcaatgaaaggagata-3’ adjacent to the PAM sequence GAGAGTA:

[0047]

Chemical formula

[0048] The reaction was carried out under the following conditions: 1xCutSmart buffer 400 nM CcCas9 20 nM DNA 2 μM crRNA 2 μM tracrRNA The incubation time was 30 minutes and the reaction temperature was 45°C. The experimental results confirmed the PAM sequence of CcCas9 as NNNNGNA-3’. The most conserved amino acid was G at the 5th position (see Figure 4). The present invention includes the following aspects in a non-limiting manner. [Aspect 1] Use of a protein that forms a double-strand break in the DNA molecule, which is located immediately before the nucleotide sequence 5'-NNNNGNA-3' in the DNA molecule, and that comprises the amino acid sequence of SEQ ID NO: 1 or is at least 95% identical to the amino acid sequence of SEQ ID NO: 1 and differs from SEQ ID NO: 1 only in non-conservative amino acid residues. [Aspect 2] Use of the protein according to aspect 1, characterized in that a double-strand break is formed in the DNA molecule at a temperature of 37°C to 65°C. [Aspect 3] Use of the protein according to aspect 1, wherein the protein comprises the amino acid sequence of SEQ ID NO: 1. [Aspect 4] A method for generating a double-strand break in the genomic DNA sequence of a unicellular or multicellular organism that is directly adjacent to the sequence 5'-NNNNGNA-3', the method comprising introducing into at least one cell of the organism an effective amount of: a) a protein comprising the amino acid sequence of SEQ ID NO: 1, or a nucleic acid encoding a protein comprising the amino acid sequence of SEQ ID NO: 1, and b) a guide RNA comprising a sequence that forms a double strand with the nucleotide sequence of the genomic DNA region of the organism that is immediately adjacent to the nucleotide sequence 5'-NNNNGNA-3' and that interacts with the protein after double-strand formation, or a DNA sequence encoding the guide RNA. Here, the interaction of the protein with the guide RNA and the nucleotide sequence 5'-NNNGNA-3' results in the formation of a double-strand break in the genomic DNA sequence that is directly adjacent to the sequence 5'-NNNNGNA-3'. The method. [Aspect 5] The method according to aspect 4, further comprising introducing an exogenous DNA sequence simultaneously with the guide RNA.

Examples

[0049] The following exemplary embodiments of the method are provided for the purpose of disclosing the characteristics of the invention and are not to be construed as limiting the scope of the invention in any way. Example 1. Testing the activity of CcCas9 protein in cutting various DNA targets.

[0050] To check the ability of CcCas9 to recognize various DNA sequences adjacent to the NNNNGNA3' sequence, an experiment on in vitro cutting of DNA targets derived from the human grin2b gene sequence (see Table 1 below) was conducted.

[0051] Table 1. DNA targets isolated from the human grin2b gene

[0052] [Table 1]

[0053] Under conditions similar to those of the above experiment, an in vitro DNA cutting reaction was carried out. As the DNA target, a human grin2b gene fragment with a size of about 500 bp was used:

[0054] [Chemical formula]

[0055] The experimental results show that CcCas9 in the complex containing guide RNA can recognize various DNA targets containing the PAM sequence NNNNGNA (Figure 5). In the case of some targets, CcCas9 tolerates substitutions at the 7th position of the PAM sequence.

[0056] Example 2. Temperature range of CcCas9 activity. To determine the temperature range of the CcCas9 protein, experiments on in vitro cutting of DNA targets were carried out under different temperature conditions.

[0057] For this purpose, target DNA adjacent to the PAM sequence GAGAGTA was subjected to cutting by the CcCas9 effector complex containing the corresponding guide RNA at different temperatures (Figure 6).

[0058] The CcCas9 protein was found to have a broad temperature range of activity. While the maximum nuclease activity is achieved at a temperature of 45°C, the protein is fully active in the range of 37°C to 55°C. Thus, CcCas9 in a complex containing guide RNA is a novel tool for cutting (formation of double-strand breaks) in DNA molecules restricted to the sequence 5'-NNNNGNA-3' with a temperature range of 37°C to 55°C. The scheme of the complex of crRNA and tracer RNA (tracrRNA) that together form the guide RNA with the target DNA is shown in Fig. 7.

[0059] Example 3. Cas9 proteins derived from related organisms belonging to the genus Clostridium. So far, only one type II CRISPR Cas system has been found in the genus Clostridium, which is the Cas9 CRISPR Cas system derived from Clostridium perfringens (Maikova A et al., New Insights Into Functions and Possible Applications of Clostridium difficile CRISPR-Cas System. Front Microbiol. 2018 Jul 31;9:1740).

[0060] The Cas9 protein derived from the bacterium Clostridium perfringens is 36% identical to the CcCas9 protein (calculated the degree of identity using BLASTp software with default parameters). The Cas9 protein derived from Staphylococcus aureus of comparable size is 28% identical to CcCas9 (BLASTp, default parameters).

[0061] Thus, the CcCas9 protein is significantly different in the amino acid sequence from other Cas9 proteins studied so far, including those found in related organisms.

[0062] One skilled in the art of genetic engineering would recognize that the CcCas9 protein sequence variants obtained and characterized by the present applicants can be modified without changing the function of the protein itself (e.g., by site-directed mutagenesis of amino acid residues that do not directly affect functional activity (Sambrook et al., Molecular Cloning: A Laboratory Manual, (1989), CSH Press, pp.15.3-15.108)). In particular, one skilled in the art would recognize that non-conserved amino acid residues can be modified without affecting residues involved in protein functionality (determining protein function or structure). Examples of such modifications include substitution of non-conserved amino acid residues with homologous ones. Some regions containing non-conserved amino acid residues are shown in FIG. 8. In some embodiments of the present invention, a protein comprising an amino acid sequence that is at least 95% identical to the amino acid sequence of SEQ ID NO: 1 and differs from SEQ ID NO: 1 only at non-conserved amino acid residues is used to form a double-strand break immediately preceding the nucleotide sequence 5'-NNNNGNA-3' in a DNA molecule. A homologous protein can be obtained by mutagenizing the corresponding nucleic acid molecule (e.g., by site-directed or PCR-mediated mutagenesis), and then the encoded modified Cas9 protein can be tested for retention of its function according to the functional analysis described herein.

[0063] Example 4. The CcCas9 system described in the present invention may be used in combination with a guide RNA to modify the genomic DNA sequence of a multicellular organism including eukaryotes. To introduce the CcCas9 system in a complex containing the guide RNA into the cells of this organism (into all cells or into some cells), various approaches known to those skilled in the art may be applied. For example, methods for delivering the CRISPR-Cas9 system into the cells of an organism are disclosed in information sources (Liu C et al., Delivery strategies of the CRISPR-Cas9 gene-editing system for therapeutic applications. J Control Release. 2017 Nov 28;266:17-26; Lino CA et al., Delivering CRISPR: a review of the challenges and approaches. Drug Deliv. 2018 Nov;25(1):1234-1257) and in information sources further disclosed within these information sources.

[0064] For effective expression of the CcCas9 nuclease in eukaryotic cells, it may be desirable to optimize the codons with respect to the amino acid sequence of the CcCas9 protein by methods known to those skilled in the art (e.g., the IDT codon optimization tool).

[0065] For effective activity of the CcCas9 nuclease in eukaryotic cells, it is necessary to translocate the protein into the nucleus of the eukaryotic cell. This may be done by using a nuclear localization signal derived from the SV40 T antigen (Lanford et al., Cell, 1986, 46:575-582), which is linked to the CcCas9 sequence either through a spacer sequence as described in Shen B et al. ”Generation of gene-modified mice via Cas9 / RNA-mediated gene targeting”, Cell Res. 2013 May;23(5):720-3 or without including a spacer sequence. Thus, the complete amino acid sequence of the nuclease to be transported into the nucleus of the eukaryotic cell would be: MAPKKKRKVGIHGVPAA-CcCas9-KRPAATKKAGQAKKKK (hereinafter referred to as CcCas9 NLS). The protein containing the above amino acid sequence may be delivered using at least two approaches.

[0066] Gene delivery is achieved by generating a plasmid that holds the CcCas9 NLS gene under the control of a promoter (e.g., CMV promoter) and a sequence encoding the guide RNA under the control of the U6 promoter. As a DNA target, a DNA sequence flanked by 5’-NNNNGNA-3’ is used, which is for example that of the human grin2b gene:

[0067]

Chemical formula

[0068] Thus, the crRNA expression cassette would be as follows:

[0069]

Chemical formula

[0070] The bold text indicates the U6 promoter sequence, followed by the sequence necessary for target DNA recognition. On the other hand, the direct repeat sequences are emphasized in capital letters. The tracer RNA expression cassette is as follows:

[0071]

Chemical formula

[0072] The bold text indicates the U6 promoter sequence, followed by the sequence encoding the tracer RNA. Purify the plasmid DNA and transfect it into human HEK293 cells using Lipofectamine 2000 reagent (Thermo Fisher Scientific). Incubate the cells for 72 hours, and then extract the genomic DNA from them using a genomic DNA purification column (Thermo Fisher Scientific). Analyze the target DNA site by sequencing on the Illumina platform to determine the number of insertions / deletions in the DNA that occurred at the target site due to directional double-strand cleavage and subsequent repair.

[0073] For example, regarding the above-mentioned grin2b gene site, amplify the target fragment using primers adjacent to the putative site of cleavage introduction:

[0074]

Chemical formula

[0075] After amplification, prepare the sample according to the Ultra II DNA Library Prep Kit for the Illumina (NEB) reagent sample preparation protocol for high-throughput array determination. Then, perform sequencing on the Illumina platform with 300 cycles of direct reads. Analyze the sequencing results by bioinformatics methods. Consider the insertion or deletion of some nucleotides in the target DNA sequence as cut detection.

[0076] Perform delivery as a ribonucleic acid complex by incubating the guide RNA and recombinant CcCas9 NLS in CutSmart buffer (NEB). Produce the recombinant protein from bacterial production cells by purifying the recombinant protein by affinity chromatography (NiNTA, Qiagen) with size exclusion (Superdex200).

[0077] Mix the protein with RNA at a ratio of 1:2:2 (CcCas9 NLS:crRNA:tracrRNA), incubate the mixture at room temperature for 10 minutes, and then transfect it into cells.

[0078] Next, analyze the DNA extracted therefrom for insertions / deletions at the target DNA site (as described above). The CcCas9 nuclease derived from the bacterium Clostridium cellulolyticum characterized in the present invention has several advantages compared to previously characterized Cas9 proteins.

[0079] Unlike other known Cas nucleases required for the system to function, CcCas9 has a short 2-letter PAM. According to the authors, the short PAM GNA located 4 nucleotides away from the protospacer is sufficient for CcCas9. Furthermore, while the G at the +5 position is very important, the +7 position is less important, and the efficiency is slightly lower in the presence of A or T at the +7 position as well as in the presence of C, but in vitro hydrolysis was detected.

[0080] Most of the Cas nucleases known to date that can introduce double-strand breaks into DNA have complex multi-letter PAM sequences, and the selection of appropriate sequences for cutting is restricted. Among the Cas nucleases studied that recognize small molecule PAMs, only CcCas9 can recognize sequences restricted to GNA nucleotides.

[0081] A second advantage of CcCas9 is its small protein size (1030 a.a.r., which is 23 a.a.r. less than that of SaCas9). To date, CcCas9 is the only low molecular weight protein studied that has a two-letter PAM sequence.

[0082] A third advantage of the CcCas9 system is its broad temperature range of activity: the nuclease is active at temperatures from 37 °C to 65 °C, with an optimum at 45 °C. Although the present invention has been described with reference to the disclosed embodiments, those skilled in the art will recognize that the specific embodiments described in detail are provided for the purpose of exemplifying the invention and are not to be considered as limiting the scope of the invention in any way. It will be understood that various modifications may be made without departing from the spirit of the invention.

Claims

1. 0 to 25 nucleotides of the nucleotide sequence 5'-NNNNGNA-3' in a DNA molecule At a position 0 to 25 nucleotides upstream, an amino acid sequence for forming a double-strand break, the amino acid sequence of SEQ ID NO: 1 or SEQ ID NO: 1 An amino acid sequence that is at least 95% identical to the amino acid sequence of, and capable of forming the double-strand break, use of the protein in vitro.

2. Characterized by the formation of a double-strand break in a DNA molecule at a temperature of 37°C to 65°C Use of the protein according to Claim 1.

3. Use of the protein according to Claim 1, wherein the protein comprises the amino acid sequence of SEQ ID NO:

1.

4. In the genomic DNA sequence of a single-celled or multi-celled organism, at a position 0 to 25 nucleotides upstream of the nucleotide sequence 5'-NNNNG NA-3', a method in vitro for generating a double-strand break, The method comprising introducing into at least one cell of the organism an effective amount of: a) a protein comprising the amino acid sequence of SEQ ID NO: 1, or a nucleic acid encoding a protein comprising the amino acid sequence of SEQ ID NO: 1, and b) a guide RNA comprising a sequence that forms a double-strand with the genomic DNA at the position 0 to 25 nucleotides upstream of the nucleotide sequence 5'-NNNNGNA-3' in the genomic DNA sequence of the organism and interacts with the protein after double-strand formation, or DNA encoding the guide RNA, Including the step of introducing, Here, the interaction of the protein with the guide RNA and the genomic DNA results in the formation of a double-strand break at the position 0 to 25 nucleotides upstream of the nucleotide sequence 5'-NNNNGNA-3' in the genomic DNA sequence Of the method.

5. The method according to Claim 4, further comprising the introduction of exogenous DNA simultaneously with the guide RNA. ​ ​ ​ ​ ​ ​ ​

Citation Information

Patent Citations

  • Materials and Methods of the CRISPR-CAS System

    JP2016537028A

  • Novel CAS systems and methods of use

    WO2017222773A1