DNA cleavage means based on the Cas9 protein derived from the species of the genus Defluviimonas
The DfCas9 enzyme from Defluviimonas species 20V17, with a unique PAM sequence and manganese requirement, addresses the limitations of CRISPR-Cas systems by enabling precise double-strand breaks in diverse organisms, enhancing genomic editing capabilities.
Patent Information
- Application Number
- JP2021529804
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-11-26
- Filing Date
- 2019-11-26
- Publication Date
- 2025-06-25
- Estimated Expiration
- 2039-11-26
AI Technical Summary
Current CRISPR-Cas systems are limited by specific PAM sequences, restricting their ability to modify DNA regions in diverse organisms, necessitating the development of novel Cas9 enzymes with alternative PAM sequences for precise double-strand break formation.
Characterization of the DfCas9 enzyme from Defluviimonas species 20V17, which has a distinct PAM sequence (5'-NN(G/A)NA(C/T)N-3') and requires manganese ions for activity, allowing for double-strand breaks at specific DNA sites.
Expands the versatility of CRISPR-Cas9 systems to cut genomic DNA at a greater number of sites, enhancing the precision and applicability to various organisms, including eukaryotes, by overcoming PAM sequence constraints.
Smart Images

Figure 0007698579000013 
Figure 0007698579000014 
Figure 0007698579000015
Abstract
Description
Technical Field
[0001] The present invention relates to a novel Cas nuclease enzyme of the CRISPR-Cas system that is used in biotechnology, particularly for cutting DNA and editing the genomes of various organisms. This technology may be used in the future for gene therapy of hereditary human diseases as well as for editing the genomes of other organisms.
Background Art
[0002] Modification of DNA sequences is one of the current issues in the field of biotechnology today. Editing and modifying the genomes of eukaryotic and prokaryotic organisms, as well as manipulating DNA in vitro, require the targeted introduction of double-strand breaks in the DNA sequence.
[0003] To solve this problem, currently, the following techniques are used: artificial nuclease systems containing zinc finger-type domains, TALEN systems, and the bacterial CRISPR-Cas system. The first two techniques require the effort of optimizing the nuclease amino acid sequence for the recognition of specific DNA sequences. In contrast, in the case of the CRISPR-Cas system, the structure that recognizes the DNA target is not a protein but a small molecule guide RNA. For the cutting of a specific DNA target, it is not necessary to newly synthesize the nuclease or its gene, but it is performed by using a guide RNA complementary to the target sequence. For this reason, the CRISPR Cas system has become a suitable and efficient means for cutting various DNA sequences. This technology enables the simultaneous cutting of DNA in several regions by using guide RNAs of different sequences. Such an approach is also used for simultaneously modifying several genes in eukaryotes.
[0004] Due to its properties, the CRISPR-Cas system is a prokaryotic immune system capable of introducing cleavage very specifically within viral genetic material (Mojica F. J. M. et al. Intervening sequences of regularly spaced prokaryotic repeats derive from foreign genetic elements / / Journal of molecular evolution. 2005. Vol. 60. Issue 2. pp.174-182). The abbreviation CRISPR-Cas stands for "Clustered Regularly Interspaced Short Palindromic Repeats and CRISPR associated Genes" (Jansen R. et al. Identification of genes that are associated with DNA repeats in prokaryotes / / Molecular microbiology. 2002. Vol. 43. Issue 6. pp.1565-1575). All CRISPR-Cas systems consist of a CRISPR cassette and genes encoding diverse Cas proteins (Jansen R. et al., Molecular microbiology. 2002. Vol. 43. Issue 6. pp.1565-1575). The CRISPR cassette consists of spacers, each having a unique nucleotide sequence, and repeat palindromic repeats (Jansen R. et al., Molecular microbiology. 2002. Vol. 43. Issue 6. pp.1565-1575).Transcription of the CRISPR cassette, followed by its processing, results in the formation of guide crRNA, which, together with the Cas protein, forms an effector complex (Brouns S. J. J. et al., Small CRISPR RNAs guide antiviral defense in prokaryotes, Science, 2008, Vol. 321, Issue 5891, pp. 960-964). Through complementary pairing between the crRNA and the target DNA site called the protospacer, the Cas nuclease recognizes the DNA target and introduces a cleavage therein in a very specific manner.
[0005] CRISPR-Cas systems containing a single effector protein are grouped into six different types (types I-VI) depending on the Cas protein included in the system. The type II CRISPR-Cas9 system is characterized by a simple composition and mechanism of action, i.e., its function requires only the formation of an effector complex consisting of one Cas9 protein and two small RNAs as follows: crRNA and tracer RNA (tracrRNA). The tracer RNA forms complementary pairs with the crRNA region generated from the CRISPR repeat to form the secondary structure necessary for the guide RNA to bind to the Cas effector. The Cas9 effector protein has two nuclease domains (HNH and RuvC) that introduce cleavage within the complementary strand of the target DNA and is thus an RNA-dependent DNA endonuclease that forms a double-stranded DNA break (Deltcheva E. et al., CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III, Nature, 2011, Vol. 471, Issue 7340, p. 602).
[0006] To date, several CRISPR-Cas nucleases capable of targeting and specifically introducing double-strand breaks into DNA are known. One of the main characteristics that limits the use of the CRISPR-Cas system is the PAM sequence, which is adjacent to the 3' end of the DNA target and whose presence is necessary for the accurate recognition of DNA by the Cas9 nuclease. Diverse CRISPR-Cas proteins have different PAM sequences, thus restricting the potential for the use of nucleases in any DNA region. To enable the modification of any DNA region both in vitro and in the biological genome, the use of CRISPR-Cas proteins with novel diverse PAM sequences is required. Modification of the eukaryotic genome also requires the use of small-sized nucleases to provide AAV-mediated delivery of the CRISPR-Cas system into cells.
[0007] Although many techniques for cutting DNA and modifying genomic DNA sequences are known, there remains a need for new and effective means for modifying DNA in diverse organisms and at precisely defined sites within the DNA sequence. The present invention provides some of the characteristics necessary to solve this problem.
Prior Art Documents
Non-Patent Documents
[0008]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Non-Patent Document 4
Summary of the Invention
Problems to be Solved by the Invention
[0009] An object of the present invention is to provide a novel means for modifying genomic DNA sequences of single-celled or multicellular organisms using the CRISPR-Cas9 system. Current systems are limited in use due to the specific PAM sequences that must be present at the 3' end of the DNA region to be modified. Searching for novel Cas9 enzymes with other PAM sequences will expand the range of available means for forming double-strand breaks at precisely the desired specific sites in DNA molecules of diverse organisms. To solve this problem, the authors characterized the type II CRISPR nuclease DfCas9, which had been previously predicted for the genus Defluviimonas species 20V17, and it is also possible to introduce directional modifications into the genomes of both the above and other organisms using this enzyme. The present invention is characterized by the following essential features: (a) a short PAM sequence different from other known PAM sequences; (b) a relatively small size of DfCas9 characterized as having 1079 amino acid residues (a.a.r.).
Means for Solving the Problems
[0010] The above problem is solved by using a protein comprising the amino acid sequence of SEQ ID NO: 1 or an amino acid sequence that is at least 95% identical to the amino acid sequence of SEQ ID NO: 1 and differs only at non-conserved amino acid residues, to form a double-strand break in the DNA molecule immediately preceding the nucleotide sequence 5'-NN(G / A)NA(C / T)N-3' in the DNA molecule. In some embodiments of the present invention, this use is characterized by the formation of a double-strand break in the DNA molecule at a temperature of 35°C to 37°C and in the presence of Mn2+ ions. In a preferred embodiment of the present invention, this use is characterized by an Mn2+ ion concentration higher than 5 mM.
[0011] The above problem is a method for generating double-strand breaks in the genomic DNA sequence of a single-celled or multi-celled organism immediately adjacent to the sequence 5'-NN(G / A)NA(C / T)N-3', comprising introducing into at least one cell of the organism an effective amount of: a) a protein comprising the amino acid sequence of SEQ ID NO: 1, or a nucleic acid encoding a protein comprising the amino acid sequence of SEQ ID NO: 1, and b) a guide RNA comprising a sequence that forms a double strand with the nucleotide sequence of the genomic DNA region of the organism immediately adjacent to the nucleotide sequence 5'-NN(G / A)NA(C / T)N-3' and interacts with the protein after double-strand formation, or a DNA sequence encoding the guide RNA, wherein the interaction between the protein, the guide RNA, and the nucleotide sequence 5'-NN(G / A)NA(C / T)N-3' results in the formation of a double-strand break in the genomic DNA sequence immediately adjacent to the sequence 5'-NN(G / A)NA(C / T)N-3', is further solved by using the above method. In some embodiments of the invention, the method is characterized by further comprising the introduction of an exogenous DNA sequence simultaneously with the guide RNA.
[0012] A mixture of crRNA and tracer RNA (tracrRNA) that can form a complex with the target DNA region and the DfCas9 protein may be used as the guide RNA. In a preferred embodiment of the invention, a hybrid RNA constructed based on crRNA and tracer RNA may be used as the guide RNA. Methods for constructing hybrid guide RNAs are known to those skilled in the art (Hsu PD et al., DNA targeting specificity of RNA-guided Cas9 nucleases. Nat Biotechnol. 2013 Sep;31(9):827-32). One approach for constructing hybrid RNA is disclosed in the following examples.
[0013] The present invention may be used both for in vitro cutting of target DNA and for modifying the genomes of several organisms. The genome may be modified either directly, i.e., by cutting the genome at the corresponding site, and by inserting exogenous DNA sequences through homologous repair.
[0014] Any region of double-stranded or single-stranded DNA derived from the genome of an organism other than for use in administration (or a composition of such regions within themselves and containing other DNA fragments) may be used as an exogenous DNA sequence, where said region (or composition of regions) is intended to be induced by DfCas9 nuclease and incorporated into the site of double-strand break in the target DNA. In some embodiments of the present invention, a double-stranded DNA region derived from the genome of an organism, which is used for the introduction of DfCas9 protein but is further modified by mutation (nucleotide substitution) and by insertion or deletion of one or more nucleotides, may be used as an exogenous DNA sequence.
[0015] The technical result of the present invention is to increase the versatility of the available CRISPR-Cas9 system and enable the use of Cas9 nuclease for cutting genomic or plasmid DNA at a greater number of specific sites and under specific conditions.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Mode for Carrying Out the Invention
[0017] When used in the description of the present invention, the terms "includes" and "including" should be construed to mean "including among other things". The terms are not intended to be construed as "consisting only of". Unless otherwise defined, technical and scientific terms in this application have the typical meanings generally recognized in scientific and technical literature.
[0018] As used herein, the term "percent homology of two sequences" is equivalent to the term "percent identity of two sequences". Sequence identity is determined based on a reference sequence. Algorithms for sequence analysis are known in the art, such as BLAST described in Altschul et al., J. Mol. Biol., 215, pp. 403-10 (1990). For the purposes of the present invention, to determine the level of identity and similarity between nucleotide and amino acid sequences, comparison of nucleotide and amino acid sequences may be used, which is performed with standard parameters and using gapped alignment by the BLAST software package provided by the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / blast). The percent identity of two sequences is determined by the number of identical amino acid positions in these two sequences, taking into account the number of gaps inserted for optimal comparison of the two sequences by alignment and the length of each gap. The percent identity equals the number of amino acids that are identical at a given position, taking into account the sequence alignment, divided by the total number of positions, and multiplied by 100.
[0019] The term "specifically hybridizes" refers to the association between two single-stranded nucleic acid molecules or sufficiently complementary sequences, which enables such hybridization under pre-determined conditions typically used in the art.
[0020] The phrase "double-stranded break located right before in the nucleotide PAM sequence" means that the double-stranded break in the target DNA sequence is made at a distance of 0 to 25 nucleotides upstream of the nucleotide PAM sequence.
[0021] The exogenous DNA sequence introduced simultaneously with the guide RNA is intended to refer to a DNA sequence specifically prepared for the specific modification of double-stranded target DNA at the cleavage site determined by the specificity of the guide RNA. Such modifications may be, for example, insertions or deletions of specific nucleotides at the cleavage site in the target DNA. The exogenous DNA may be either a DNA region from a different organism or a DNA region from the same organism as that of the target DNA.
[0022] A protein containing a specific amino acid sequence is intended to refer to a protein having the amino acid sequence and an amino acid sequence composed of other sequences that can be linked to the amino acid sequence by peptide bonds. Examples of other sequences may be a nuclear localization signal (NLS) or other sequences that provide an increase in the functionality of the amino acid sequence.
[0023] The exogenous DNA sequence introduced simultaneously with the guide RNA is intended to refer to a DNA sequence specifically prepared for the specific modification of double-stranded target DNA at the site of cleavage determined by the specificity of the guide RNA. Such modifications may be, for example, insertions or deletions of specific nucleotides at the cleavage site in the target DNA. The exogenous DNA may be either a DNA region from a different organism or a DNA region from the same organism as that of the target DNA.
[0024] The effective amounts of the protein and RNA introduced into the cell are intended to refer to the amounts of the protein and RNA such that, when introduced into the cell, a functional complex, i.e., a complex that can specifically bind to the target DNA and cause a double-strand break at the site determined by the guide RNA and the PAM sequence on the DNA, can be formed. The efficiency of this process can be evaluated by analyzing the target DNA isolated from the cell using conventional techniques known to those skilled in the art.
[0025] Proteins and RNAs may be delivered into cells by a variety of techniques. For example, a protein may be delivered as a DNA plasmid encoding the gene of this protein, as mRNA that is translated into this protein in the cytoplasm, or as a ribonucleoprotein complex containing this protein and guide RNA. Delivery may be performed by a variety of techniques known to those skilled in the art.
[0026] The nucleic acid encoding the components of the system may be introduced into cells directly or indirectly, intracellularly, by transfection or transformation of cells by methods known to those skilled in the art, by use of recombinant viruses, or by manipulation of cells, such as DNA microinjection, etc.
[0027] A ribonucleic complex consisting of a nuclease, guide RNA, and exogenous DNA (if necessary) may be delivered by transfecting the complex into cells or by mechanically introducing the complex into cells, such as by microinjection.
[0028] The nucleic acid molecule encoding the protein to be introduced into cells may be integrated into the chromosome or may be DNA replicated extrachromosomally. In some embodiments, in order to ensure efficient expression of the protein gene in the introduced DNA, since the frequency of occurrence of synonymous codons is uneven in the coding regions of various biological genomes, it is necessary to modify the sequence of said DNA according to the cell type so as to optimize the codons for expression. Codon optimization is necessary to increase expression in animal, plant, fungal, or microbial cells.
[0029] For a protein having a sequence that is at least 95% identical to the amino acid sequence of SEQ ID NO: 1 to function in a eukaryotic cell, it is necessary for this protein to ultimately reach the nucleus of this cell. Thus, in some embodiments of the present invention, a protein having a sequence that is at least 95% identical to the amino acid sequence of SEQ ID NO: 1 and that is further modified at one or both ends by the addition of one or more nuclear localization signals is used to form a double-strand break in target DNA. For example, a nuclear localization signal derived from the SV40 virus may be used. To provide efficient delivery to the nucleus, for example, the spacer sequence may be used to separate the nuclear localization signal from the major protein sequence as described in Shen B et al., "Generation of gene-modified mice via Cas9 / RNA-mediated gene targeting", Cell Res. 2013 May;23(5):720-3. Further, in other embodiments, different nuclear localization signals, or alternative methods for delivering the protein into the cell nucleus may be used.
[0030] The present invention includes the use of a protein derived from the Defluviimonas sp. 20V17 organism that is homologous to a previously characterized Cas9 protein for introducing a double-strand break at a precisely specified position within a DNA molecule. The use of CRISPR nucleases for introducing targeted modifications into the genome has several advantages. First, the specificity of the activity of the system is determined by the crRNA sequence, which allows for the use of one type of nuclease for all target loci. Second, the technology allows for the simultaneous delivery into the cell of several guide RNAs that are complementary to different gene targets, thereby allowing for the simultaneous modification of several genes.
[0031] For the biochemical characterization of the Cas9 protein derived from the Defluviimonas sp. 20V17 bacterium, the CRISPR locus encoding the main system components (DfCas9, cas1, cas2 protein genes, as well as the CRISPR cassette and guide RNA) was cloned into the single-copy bacterial vector pACYC184. The effector ribonucleic acid complex consisting of Cas9 and the crRNA / tracrRNA duplex requires the presence of a PAM (protospacer adjacent motif) on the DNA target for DNA recognition and subsequent hydrolysis, in addition to crRNA spacer-protospacer complementarity (Mojica F. J. M. et al. Short motif sequences determine the targets of the prokaryotic CRISPR defence system / / Microbiology. 2009. Vol. 155. Issue 3. pp.733-740). The PAM is a strictly defined sequence of several nucleotides located within type II systems, adjacent to or a few nucleotides away from the 3' end of the protospacer on the non-target strand. In the absence of the PAM, hydrolysis of DNA binding with double-strand break formation does not occur. The requirement for the presence of the PAM sequence on the target increases recognition specificity but simultaneously imposes constraints on the selection of the target DNA region for introducing cleavage.
[0032] To determine the sequence of the guide RNA of the CRISPR-Cas9 system, RNA sequencing of Escherichia coli (E. coli) DH5 alpha bacteria carrying the generated DfCas9_pacyc184 construct was performed. Sequencing showed that the CRISPR cassette of the system was actively transcribed, as was the tracer RNA (Figure 1). Analysis of the crRNA and tracrRNA sequences enabled the possibility of considering that they could form secondary structures that are presumably recognized by the DfCas9 nuclease.
[0033] Furthermore, the authors determined the PAM sequence of the DfCas9 protein using bacterial PAM screening. To determine the PAM sequence of the DfCas9 protein, Escherichia coli DH5 alpha cells carrying the DfCas9_pacyc184 plasmid were transformed with a plasmid library containing the spacer sequence 5’-TAGACCTTCGGGATCATGTCGATCATGATC-3’ of the DfCas9-based CRISPR cassette, with a random 7-character sequence adjacent to either the 5’ or 3’ end. Plasmids carrying sequences corresponding to the DfCas9-based PAM sequence were subjected to degradation under the action of a functional CRISPR-Cas system, while the remaining library plasmids were effectively transformed into the cells and made resistant to the antibiotic ampicillin. After transformation and incubation of the cells on plates containing the antibiotic, the colonies were washed off the agar surface and DNA was extracted from them using the Qiagen Plasmid Purification Midi Kit. From the plasmid isolation pool, the region containing the randomized PAM sequence was amplified by PCR and then subjected to high-throughput sequencing on the Illumina platform. The resulting reads were analyzed by comparing the transformation efficiency of plasmids containing unique PAMs included in the library into cells carrying DfCas9_pacyc184 or into control cells carrying the empty vector pacyc184. The results were analyzed using bioinformatics methods. As a result, it was possible to identify the DfCas9-based PAM, which is the 3-character sequence 5’-NN(G / A)NA(C / T)N-3’ (Figure 2).
[0034] To develop a system for cutting DNA in vivo and in vitro, it is necessary to obtain all the components that are part of the effector DfCas9 complex as follows: guide RNA and DfCas9 nuclease. Determination of the guide RNA sequence by RNA sequencing made it possible to synthesize the crRNA and tracrRNA molecules in vitro. Synthesis was performed using the NEB HiScribe T7 RNA Synthesis Kit.
[0035] To cut a DNA target restricted by the 3'-end sequence NN(G / A)NA(C / T)N, a guide RNA with the following sequence was used:
[0036]
Chemical formula
[0037] The first 20 nucleotides contained in the crRNA sequence base pair complementarily with the corresponding sequence on the DNA target and provide its specific recognition by the DfCas9 nuclease. If it is desired to cut different DNA targets, this 20-character spacer sequence is modified (Figure 3).
[0038] To obtain the recombinant DfCas9 protein, its gene was cloned into the plasmid pET21a. Escherichia coli Rosetta cells were transformed with the resulting plasmid DfCas9_pET21a. Cells carrying the plasmid were grown to an optical density of OD600 = 0.6, and then the expression of the DfCas9 gene was induced by adding IPTG to a concentration of 1 mM. After incubating the cells at 25 °C for 4 hours, they were lysed. The recombinant protein was purified in two steps as follows: by affinity chromatography (Ni-NTA) and by protein size exclusion on a Superdex 200 column. The resulting protein was concentrated using an Amicon 30 kDa filter. Subsequently, the protein was frozen at -80 °C and used for in vitro reactions.
[0039] As the target DNA, a double-stranded DNA fragment with a length of 378 base pairs (bp) was used, and the DNA had a protospacer sequence at a distance of approximately 30 bp from the 3'-end restricted by the corresponding PAM sequence: NN(G / A)NA(C / T)N, as determined in experiments on bacteria.
[0040]
Chemical formula
[0041] The activity of the nuclease domain of the Cas9 protein requires the presence of divalent ions. In this regard, in vitro reactions regarding the cutting of DNA targets were carried out in the presence of magnesium, manganese, calcium and zinc salts. The in vitro reaction for cutting the DNA target was carried out under the following conditions: 1x Tris-HCl buffer, pH = 7.9 (25 °C) 400 nM DfCas9 20 nM DNA target 2 μM crRNA 2 μM tracrRNA 1 mM salt of the corresponding ion The total reaction volume was 20 μl, the reaction time was 30 minutes, and the incubation temperature was 37 °C.
[0042] The reaction products were applied onto a 1.5% agarose gel and subjected to electrophoresis. If the cutting of the DNA target was successful, the fragment was split into two parts, one of which (length approximately 325 bp) formed a distinguishable band on the gel (Figure 4). The experimental results showed that the DfCas9 protein requires manganese ions rather than magnesium ions for its own activity, which is an unusual property regarding the Cas9 nucleases characterized to date.
[0043] To confirm the significance of different DfCas9 PAM sequences, experiments regarding the in vitro cutting of DNA targets were carried out, which were similar to those previously used but contained point (single nucleotide) mutations in the PAM sequence 5'-AAAAACG-3' (Figure 5).
[0044] The experimental results confirmed that the DfCas9 enzyme is restricted by the 3'-end PAM sequence NN(G / A)NA(C / T)N and can introduce double-strand breaks into the DNA sequence. Substitutions at positions 3, 5, and 6 of the PAM are very important for the DfCas9 effector complex and prevent effective cutting of the target DNA.
[0045] Furthermore, experiments were conducted to find the temperature optimal for DfCas9 activity: it was found to be 35 - 37 °C, which gives rise to the prospect of using DfCas9 in human cells. The present invention includes the following aspects non - restrictively. [Aspect 1] Use of a protein that contains the amino acid sequence of SEQ ID NO: 1, or is at least 95% identical to the amino acid sequence of SEQ ID NO: 1 and differs from SEQ ID NO: 1 only by non - conservative amino acid residues, for forming a double - strand break in a DNA molecule, located immediately before the nucleotide sequence 5’ - NN(G / A)NA(C / T)N - 3’ in the DNA molecule. [Aspect 2] The use according to Aspect 1, characterized in that a double - strand break is formed in the DNA molecule at a temperature of 35 °C to 37 °C and in the presence of Mn2+ ions. [Aspect 3] Use of the protein according to Aspect 1, wherein the protein contains the amino acid sequence of SEQ ID NO: 1. [Aspect 4] A method for generating a double - strand break in the genomic DNA sequence of a single - celled or multi - celled organism directly adjacent to the sequence 5’ - NN(G / A)NA(C / T)N - 3’, comprising introducing into at least one cell of the organism an effective amount of: a) a protein containing the amino acid sequence of SEQ ID NO: 1, or a nucleic acid encoding a protein containing the amino acid sequence of SEQ ID NO: 1, and b) a guide RNA containing a sequence that forms a double - strand with the nucleotide sequence of the genomic DNA region of the organism directly adjacent to the nucleotide sequence 5’ - NN(G / A)NA(C / T)N - 3’ and interacts with the protein after double - strand formation, or a DNA sequence encoding the guide RNA. Here, the interaction between the protein, the guide RNA, and the nucleotide sequence 5’ - NN(G / A)NA(C / T)N - 3’ results in the formation of a double - strand break in the genomic DNA sequence directly adjacent to the sequence 5’ - NN(G / A)NA(C / T)N - 3’. The said method. [Aspect 5] The method according to embodiment 4, further comprising the introduction of an exogenous DNA sequence, simultaneously with the guide RNA.
Example
[0046] The following exemplary embodiments of the method are provided for the purpose of disclosing the characteristics of the present invention and are not to be construed as limiting the scope of the present invention in any sense. Example 1. Testing the activity of DfCas9 protein in cutting various DNA targets.
[0047] To test the ability of DfCas9 to recognize different DNA sequences flanked by the motif 5’-NN(G / A)NA(C / T)N-3’, experiments were conducted on in vitro cutting of DNA targets (see Table 1) derived from the human grin2b gene sequence.
[0048] Table 1. DNA targets isolated from the human grin2b gene
[0049]
Table 1
[0050] Under conditions similar to the previous experiments, in vitro DNA cutting reactions were performed. Guide crRNAs were synthesized for each of the target sequences. As the DNA target, a human grin2b gene fragment of approximately 760 bp in size was used:
[0051]
Chemical formula
[0052] A DNA region corresponding to the PAM sequence 5’-NN(G / A)NA(C / T)N-3’ adjacent to the 3’ end was selected as the DNA target. Among the six selected sequences, only two were effectively cut by the DfCas9 protein: DNA fragments of appropriate size cut by the effector complex were seen on the gel (Figure 6). The selectivity in the cutting of different targets can be explained by the different efficiencies of DfCas9 recognition of different secondary DNA structures or other reasons. The DfCas9 selectivity in the cutting of different targets can increase the specificity of the protein when cutting the genome of eukaryotic cells.
[0053] Therefore, DfCas9 combined with guide RNA is a novel means of double-stranded DNA cleavage restricted to the sequence 5’-NN(G / A)NA(C / T)N-3’ with a characteristic activity temperature range of 35 °C to 37 °C.
[0054] Example 2. Effect of Mn2+ ion concentration on nuclease functionality. To test the effect of Mn2+ ions on DfCas9 nuclease activity, an in vitro experiment was conducted on the cleavage of double-stranded DNA fragments containing a target DNA sequence restricted to the PAM sequence with the consensus sequence 3’-NN(G / A)NA(C / T)-5’. The reaction was carried out using 1xCutSmart buffer (NEB), target DNA at a concentration of 20 nM, and trRNA / crRNA at a concentration of 2 μM.
[0055]
Chemical formula
[0056] As the DNA target, the inventors
[0057] DNA target, the inventors
[0058]
Chemical formula
[0059] was used. The reaction was carried out using tracrRNA:
[0060]
Chem.
[0061] was used. Figure 7 shows that an increase in the concentration of MnCl2 from 5 mM to 10 mM results in more efficient cutting of the DNA target, while a further increase in the divalent ion concentration has no effect on the reaction efficiency. Thus, effective DfCas9 activity requires the presence of manganese ions at a concentration of 10 mM or higher.
[0062] Example 3. Use of hybrid guide RNA for cutting DNA targets. sgRNA is one type of guide RNA and is a fused tracrRNA (trans-activating crRNA) and crRNA. To select the optimal sgRNA, we constructed three variants of this sequence with different lengths of the tracrRNA-crRNA duplex. The RNAs were synthesized in vitro and experiments were performed on them for the cutting of DNA targets (Figure 8 shows the DNA cutting reaction by DfCas9 using various sgRNA variants).
[0063] The selected sgRNA was as effective as the native tracrRNA and crRNA sequences, and more than 50% of the DNA target was cut. When modifying the sequence that directly pairs with the DNA target, this variant of sgRNA can be used to cut any other target DNA.
[0064] The following RNA sequences were used as hybrid RNAs:
[0065]
Chem.
[0066] The "Tai" character indicates a 20-nucleotide sequence that provides pairing with the target DNA (the variable part of the sgRNA). Experiments were conducted using a control sample without RNA and a positive control of target cleavage using crRNA + trRNA.
[0067] The reaction was carried out under the following conditions: the concentration of the DNA sequence containing PAM (AAAACG) was 20 nM, the protein concentration was 400 nM, the RNA concentration was 2 μM; the incubation time was 30 minutes, and the incubation temperature was 37 °C.
[0068] After modifying the sequence that directly pairs with the DNA target, this hybrid RNA mutant may be used to cut any other target DNA. Example 4. Cas9 protein derived from related organisms belonging to the genus Defluviimonas.
[0069] To date, the CRISPR-Cas9 enzyme has not been characterized in the genus Defluviimonas. The Cas9 protein derived from Staphylococcus aureus of comparable size is 21% identical to DfCas9, and the Cas9 derived from Campylobacter jejuni is 28% identical to DfCas9 (calculated the degree of identity using BLASTp software with default parameters).
[0070] Therefore, the amino acid sequence of the DfCas9 protein is significantly different from other Cas9 proteins studied to date. One skilled in the art of genetic engineering would recognize that the DfCas9 protein sequence variants obtained and characterized by the present applicants can be modified without changing the function of the protein itself (e.g., by site-directed mutagenesis of amino acid residues that do not directly affect functional activity (Sambrook et al., Molecular Cloning: A Laboratory Manual, (1989), CSH Press, pp. 15.3-15.108)). In particular, one skilled in the art would recognize that non-conserved amino acid residues can be modified without affecting residues involved in protein functionality (determining protein function or structure). Examples of such modifications include substitution of non-conserved amino acid residues with homologous ones. Some regions containing non-conserved amino acid residues are shown in Figure 9. In some embodiments of the present invention, a protein comprising an amino acid sequence that is at least 95% identical to the amino acid sequence of SEQ ID NO: 1 and differs from SEQ ID NO: 1 only in non-conserved amino acid residues can be used to form a double-strand break immediately preceding the nucleotide sequence 5'-NN(G / A)NA(C / T)N-3' in a DNA molecule. A homologous protein can be obtained by mutagenizing the corresponding nucleic acid molecule (e.g., by site-directed or PCR-mediated mutagenesis), and then the encoded modified Cas9 protein can be tested for retention of its function according to the functional analysis described herein.
[0071] Example 5. The DfCas9 system described in the present invention may be used in combination with a guide RNA to modify the genomic DNA sequence of a multicellular organism including eukaryotes. To introduce the DfCas9 system in a complex containing the guide RNA into the cells of this organism (in all cells or in some cells), various approaches known to those skilled in the art may be applied. For example, methods for delivering the CRISPR-Cas9 system to the cells of an organism are disclosed in information sources (Liu C et al., Delivery strategies of the CRISPR-Cas9 gene-editing system for therapeutic applications. J Control Release. 2017 Nov 28;266:17-26; Lino CA et al., Delivering CRISPR: a review of the challenges and approaches. Drug Deliv. 2018 Nov;25(1):1234-1257) and in information sources further disclosed within these information sources.
[0072] For efficient expression of the DfCas9 nuclease in eukaryotic cells, it may be desirable to optimize the codons with respect to the amino acid sequence of the DfCas9 protein by methods known to those skilled in the art (e.g., the IDT codon optimization tool).
[0073] For effective activity of DfCas9 nuclease in eukaryotic cells, it is necessary to translocate the protein into the nucleus of eukaryotic cells. This may be done by using a nuclear localization signal derived from SV40 T antigen (Lanford et al., Cell, 1986, 46:575 - 582), which is linked to the DfCas9 sequence either through the spacer sequence described in Shen B et al. ”Generation of gene - modified mice via Cas9 / RNA - mediated gene targeting”, Cell Res. 2013 May;23(5):720 - 3 or without including the spacer sequence. Thus, the complete amino acid sequence of the nuclease to be transported into the nucleus of eukaryotic cells would be: MAPKKKRKVGIHGVPAA - DfCas9 - KRPAATKKAGQAKKKK (hereinafter referred to as DfCas9 NLS). Proteins containing the above - mentioned amino acid sequence may be delivered using at least two approaches.
[0074] Gene delivery is achieved by generating a plasmid that holds the DfCas9 NLS gene under the control of a promoter (e.g., CMV promoter) and a sequence encoding guide RNA under the control of the U6 promoter. As the target DNA, a DNA sequence flanked by 3’ - NN(G / A)NA(C / T) - 5’ is used, which is, for example, the sequence of the human grin2b gene:
[0075]
Chemical formula
[0076] Thus, the crRNA expression cassette would be as follows:
[0077]
Chemical formula
[0078] The bold text indicates the U6 promoter sequence, followed by the sequence necessary for target DNA recognition. On the other hand, the direct repeat sequences are emphasized in capital letters. The tracer RNA expression cassette is as follows:
[0079]
Chemical formula
[0080] The bold text indicates the U6 promoter sequence, followed by the sequence encoding the tracer RNA. Purify the plasmid DNA and transfect it into human HEK293 cells using Lipofectamine 2000 reagent (Thermo Fisher Scientific). Incubate the cells for 72 hours, and then extract the genomic DNA from them using a genomic DNA purification column (Thermo Fisher Scientific). Analyze the target DNA site by sequencing on the Illumina platform to determine the number of insertions / deletions in the DNA that occurred at the target site due to directed double-strand cleavage and subsequent repair.
[0081] For example, regarding the above-mentioned grin2b gene site, amplify the target fragment using primers adjacent to the putative site of cleavage introduction:
[0082]
Chemical formula
[0083] After amplification, samples are prepared according to the Ultra II DNA Library Prep Kit with respect to the Illumina (NEB) reagent sample preparation protocol for high-throughput array determination. Next, sequencing is performed on the Illumina platform with 300 cycles of direct reads. The sequencing results are analyzed by bioinformatics methods. Insertions or deletions of some nucleotides in the target DNA sequence are regarded as cut detections.
[0084] Delivery as a ribonucleic acid complex is performed by incubating the guide RNA and recombinant DfCas9 NLS in CutSmart buffer (NEB). Recombinant protein is produced from bacterial production cells by purifying the recombinant protein by affinity chromatography (NiNTA, Qiagen) with size exclusion (Superdex200).
[0085] The protein is mixed with RNA at a ratio of 1:2:2 (DfCas9 NLS:crRNA:tracrRNA), the mixture is incubated at room temperature for 10 minutes, and then transfected into cells.
[0086] Next, the DNA extracted therefrom is analyzed for insertions / deletions at the target DNA site (as described above). The DfCas9 nuclease characterized in the present invention from the deep-sea bacterium Defluviimonas sp. 20V17 has several advantages compared to previously characterized Cas9 proteins. For example, DfCas9 has a relatively simple three-letter PAM, different from other known Cas nucleases, and the PAM is required for the system to function. Most of the currently known Cas nucleases capable of introducing double-strand breaks in DNA have complex multi-letter PAMs, which limit the selection of appropriate sequences for cutting. Among the Cas nucleases that recognize short PAMs studied to date, only DfCas9 can recognize sequences restricted to 5’-NN(G / A)NA(C / T)N-3’ nucleotides.
[0087] The second advantage of DfCas9 is its relatively small protein size (1079 a.a.r.). To date, the enzyme is one of the few small-sized proteins studied that has a simple PAM sequence and is active at 37°C. The unique property of DfCas9 is that it requires the presence of manganese ions to successfully cut DNA targets. This property can be used to control the activity of the Cas9 complex.
[0088] The present invention has been described with reference to the disclosed embodiments, but those skilled in the art will recognize that the specific embodiments described in detail are provided for the purpose of exemplifying the present invention and are not considered to limit the scope of the present invention in any way. It will be understood that various modifications may be made without departing from the spirit of the present invention.
Claims
1. At a position 0 to 25 nucleotides upstream of the nucleotide sequence 5'-NN(G / A)NA(C / T)N-3' in a DNA molecule, an amino acid sequence of SEQ ID NO: 1 or an amino acid sequence that is at least 95% identical to the amino acid sequence of SEQ ID NO: 1, which can form a double-strand break, for use in vitro of a protein.
2. The use according to claim 1, characterized in that a double-strand break is formed in a DNA molecule at a temperature of 35°C to 37°C and in the presence of Mn2+ ions.
3. The use of the protein according to claim 1, wherein the protein comprises the amino acid sequence of SEQ ID NO:
1.
4. An in vitro method for generating a double-strand break at a position 0 to 25 nucleotides upstream of the nucleotide sequence 5'-NN(G / A )NA(C / T)N-3' in the genomic DNA sequence of a cell or a multicellular organism, wherein the method comprises introducing into at least one cell of the organism an effective amount of: a) a protein comprising the amino acid sequence of SEQ ID NO: 1, or a nucleic acid encoding a protein comprising the amino acid sequence of SEQ ID NO: 1, and b) a guide RNA comprising a sequence that forms a double-strand with the genomic DNA at the position 0 to 25 nucleotides upstream of the nucleotide sequence 5'-NN(G / A)NA(C / T)N-3' in the genomic DNA sequence of the organism and interacts with the protein after double-strand formation, or DNA encoding the guide RNA, wherein the interaction of the protein with the guide RNA and the genomic DNA results in the formation of a double-strand break at the position 0 to 25 nucleotides upstream of the nucleotide sequence 5'-NN(G / A)NA(C / T)N-3' in the genomic DNA sequence.
5. The method according to claim 4, further comprising introducing exogenous DNA simultaneously with the guide RNA.