Use of a Cas9 protein derived from the bacterium Pasteurella pneumotropica

The PpCas9 protein with a novel PAM sequence and small size expands the CRISPR-Cas9 system's capabilities for precise genomic editing, addressing limitations in existing systems by enabling targeted double-strand breaks and modifications across diverse organisms.

JP7708752B2Active Publication Date: 2025-07-15JOINT CO BIOCAD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2022527121
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-11-11
Filing Date
2020-07-02
Publication Date
2025-07-15
Estimated Expiration
2040-07-02

AI Technical Summary

Technical Problem

Existing CRISPR-Cas9 systems are limited by the requirement for a specific PAM sequence at the 3'-end of the DNA region, restricting their use in modifying genomic DNA sequences, necessitating the development of novel Cas9 enzymes with alternative PAM sequences to enable targeted double-strand breaks at precise sites.

Method used

Utilization of a Cas9 protein derived from Pasteurella pneumotropica (PpCas9) with a unique PAM sequence (5'-NNNN(A/G)TT-3') and a small size, allowing for targeted double-strand breaks in DNA molecules, facilitated by a guide RNA that forms a complex with the protein to introduce modifications at specific genomic sites.

Benefits of technology

Enhances the versatility of CRISPR-Cas9 systems by enabling precise DNA cleavage and modification at more specific sites, applicable to various organisms including mammalian cells, with improved efficiency and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708752000015
    Figure 0007708752000015
  • Figure 0007708752000016
    Figure 0007708752000016
  • Figure 0007708752000017
    Figure 0007708752000017
Patent Text Reader

Abstract

The present invention describes a novel bacterial nuclease from the P. pneumophila bacteria CRISPR-Cas9 system, and its use for forming strictly specific double-strand breaks in DNA molecules.This nuclease has unique properties and can be used as a tool for modifying genomic DNA sequences in unicellular organisms or multicellular organisms.Therefore, the versatility of the available CRISPR-Cas9 system is increased, and this fact allows the use of various variants of Cas9 nuclease to cut the genome or plasmid DNA of various organisms at more specific sites and / or under various conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to biotechnology, in particular, a novel enzyme, a Cas nuclease of the CRISPR-Cas system, which is used to cleave the DNA of various organisms and edit the genome. This technology can be used in the future for gene therapy of hereditary human diseases as well as genome editing of other organisms.

Background Art

[0002] DNA sequence modification is one of the topical issues in today's field of biotechnology. To edit and modify the genomes of eukaryotes and prokaryotes and manipulate DNA in vitro, it is necessary to introduce double-strand breaks targeted to the DNA sequence.

[0003] To solve this problem, the following techniques: artificial nuclease systems containing zinc finger-type domains, TALEN systems, and bacterial CRISPR-Cas systems have been used so far. The first two techniques require cumbersome optimization of the nuclease amino acid sequence to recognize specific DNA sequences. In contrast, in the case of the CRISPR-Cas system, the structure that recognizes the DNA target is a short guide RNA rather than a protein. Cleavage of a specific DNA target is performed not by de novo synthesis of the nuclease or its gene, but by using a guide RNA complementary to the target sequence. This makes the CRISPR Cas system a convenient and effective means for cleaving various DNA sequences. This technology enables simultaneous cleavage of DNA in several regions using guide RNAs of different sequences. This approach is also used to simultaneously modify several genes in eukaryotes.

[0004] Due to its properties, the CRISPR-Cas system is a prokaryotic immune system that can introduce cleavage very specifically into viral genetic material (Mojica F.J.M. et al., Intervening sequences of regularly spaced prokaryotic repeats derive from foreign genetic elements / / Journal of molecular evolution, 2005, Vol. 60, No. 2, pp. 174-182). The abbreviation CRISPR-Cas represents "clustered regularly interspaced short palindromic repeats and CRISPR-associated genes" (Jansen R. et al., Identification of genes that are associated with DNA repeats in prokaryotes / / Molecular microbiology, 2002, Vol. 43, No. 6, pp. 1565-1575). All CRISPR-Cas systems consist of a CRISPR cassette and genes encoding various Cas proteins (Jansen R. et al., Molecular microbiology, 2002, Vol. 43, No. 6, pp. 1565-1575). The CRISPR cassette consists of spacers with unique nucleotide sequences and repeated palindromic repeats (Jansen R. et al., Molecular microbiology, 2002, Vol. 43, No. 6, pp. 1565-1575). Transcription of the CRISPR cassette, followed by processing, results in the formation of guide crRNAs that form effector complexes with Cas proteins (Brouns S.J.J. et al., Small CRISPR RNAs guide antiviral defense in prokaryotes / / Science, 2008, Vol. 321, No. 5891, pp. 960-964). Through complementary pairing between the crRNA and a target DNA site called the protospacer, the Cas nuclease recognizes the DNA target and introduces cleavage very specifically into it.

[0005] CRISPR-Cas systems containing a single effector protein are classified into six different types (types I-VI) depending on the Cas protein included in the system. In 2013, it was first proposed to use the type II CRISPR-Cas9 system to edit genomic DNA in human cells (Cong L, et al., Multiplex genome engineering using CRISPR / Cas systems, Science, February 15, 2013; 339(6121):819-23). The type II CRISPR-Cas9 system is characterized by its simple composition and mechanism of activity, that is, its function requires the formation of an effector complex consisting of only one Cas9 protein and two short RNAs: crRNA and tracer RNA (tracrRNA). The tracer RNA pairs complementarily with the crRNA region derived from the CRISPR repeat to form the secondary structure necessary for the binding of the guide RNA to the Cas effector. Determining the sequence of the guide RNA is an important step in the characterization of Cas homolog species that have not been studied so far. The Cas9 effector protein is an RNA-dependent DNA endonuclease with two nuclease domains (HNH and RuvC) that introduce cleavage into the complementary strand of the target DNA, thus causing double-strand DNA cleavage (Deltcheva E. et al. CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III / / Nature, 2011, Vol. 471, No. 7340, 602).

[0006] To date, several CRISPR-Cas nucleases that can specifically introduce double-strand breaks into DNA are known. The CRISPR-Cas9 technology is one of the newest and rapidly developing technologies for introducing breaks into the DNA of various organisms from strains to human cells, and also provides in vitro applications (Song M. The CRISPR / Cas9 system: Their deliverly, in vivo and ex vivo applications and clinical development by startups., Biotechnol Prog. July 2017; 33(4): 1035-1045).

[0007] The effector ribonucleic acid complex consisting of Cas9 and the crRNA / tracrRNA duplex requires the presence of a PAM (protospacer adjacent motif) on the DNA target for DNA recognition and subsequent hydrolysis, in addition to the complementarity of the crRNA spacer-protospacer (Mojica F.J.M. et al. 2009). In type II systems, the PAM is a strictly defined sequence of several nucleotides located adjacent to or a few nucleotides away from the 3'-end of the protospacer on the non-target strand. In the absence of PAM, DNA-binding hydrolysis and subsequent double-strand break formation do not occur. The requirement for the presence of a PAM sequence on the target increases the recognition specificity but simultaneously restricts the selection of the target DNA region for introducing breaks. Therefore, the presence of the desired PAM sequence adjacent to the DNA target from the 3'-end side is a feature that limits the use of the CRISPR-Cas system at any DNA site.

[0008] Different CRISPR-Cas proteins use different specific PAM sequences for their activity. Using CRISPR-Cas proteins with various novel PAM sequences is necessary to enable modification of any DNA region both in vitro and within the genomes of organisms. Modification of eukaryotic genomes also requires the use of nucleases of small size to achieve AAV-mediated delivery of the CRISPR-Cas system into cells.

[0009] Although numerous techniques for cleaving DNA and modifying genomic DNA sequences are known, there is still a need for a new and effective means of modifying DNA at precisely specific sites in the DNA sequences of various organisms.

Summary of the Invention

Problems to be Solved by the Invention

[0010] An object of the present invention is to provide a new means for modifying the genomic DNA sequence of a single-cell or multicellular organism using the CRISPR-Cas9 system. Due to the specific PAM sequence that must be present at the 3'-end of the DNA region to be modified, the existing systems have limited use. Searching for novel Cas9 enzymes with other PAM sequences will expand the range of means available to form double-strand breaks at desired, precisely specific sites in the DNA molecules of various organisms. To solve this problem, the authors characterized the type II CRISPR nuclease PpCas9, which had been previously predicted for Pasteurella pneumotropica (P. pneumotropica), and it can be used to introduce targeted modifications into the genomes of both the above and other organisms. The present invention is characterized by essential features: (a) a short PAM sequence different from other known PAM sequences; (b) a characterized PpCas9 protein of a relatively small size, which is 1055 amino acid residues (a.a.r.).

Means for Solving the Problems

[0011] The above problem is solved by the use of a protein comprising the amino acid sequence of SEQ ID NO: 1, or an amino acid sequence that is at least 95% identical to the amino acid sequence of SEQ ID NO: 1 and differs from SEQ ID NO: 1 only at non-conserved amino acid residues, for forming a double-strand break in a DNA molecule immediately prior to the nucleotide sequence 5'-NNNN(A / G)TT-3' within the DNA molecule. In some embodiments of the present invention, this use is characterized in that the double-strand break in the DNA molecule is formed at a temperature of 35°C to 45°C. In some embodiments of the present invention, this use is characterized in that the double-strand break is formed within the genomic DNA of mammalian cells. In some embodiments of the present invention, this use is characterized in that the formation of the double-strand break in the DNA molecule results in the modification of the genomic DNA of the mammalian cells.

[0012] The above problem is further solved by providing a method for modifying the genomic DNA sequence of a cell of a unicellular or multicellular organism, the method comprising the step of introducing into the cell of the organism an effective amount of: a) a protein comprising the amino acid sequence of SEQ ID NO: 1 or a nucleic acid encoding a protein comprising the amino acid sequence of SEQ ID NO: 1, and b) a guide RNA comprising a sequence that forms a double-strand with the nucleotide sequence of a genomic DNA region of the organism immediately adjacent to the nucleotide sequence 5'-NNNN(A / G)TT-3' and interacts with the protein after double-strand formation, or a DNA sequence encoding the guide RNA, wherein the interaction of the guide RNA and the nucleotide sequence 5'-NNNN(A / G)TT-3' with the protein results in the formation of a double-strand break in the genomic DNA sequence immediately adjacent to the sequence 5'-NNNN(A / G)TT-3'.

[0013] In some embodiments of the present invention, the method further comprises the step of introducing an exogenous DNA sequence simultaneously with the guide RNA. In some embodiments of the present invention, the method is characterized in that the cell is a mammalian cell.

[0014] A mixture of a target DNA region and a crRNA and a tracer RNA (tracrRNA) that can form a complex with the PpCas9 protein may be used as a guide RNA. In a preferred embodiment of the present invention, a hybrid RNA constructed based on crRNA and tracrRNA may be used as a guide RNA. Methods for constructing hybrid guide RNAs are known to those skilled in the art (Hsu PD, et al., DNA targeting specificity of RNA-guided Cas9 nucleases. Nat Biotechnol. September 2013; 31(9): 827-832). One of the techniques for constructing hybrid RNA is disclosed in the following examples.

[0015] The present invention can be used for both in vitro cleavage of target DNA and modification of the genomes of several organisms. Genomic DNA can be modified by a direct method, namely, a step of cleaving genomic DNA at the corresponding site and a step of inserting an exogenous DNA sequence using homologous repair.

[0016] Any region of double-stranded or single-stranded DNA derived from the genome of an organism other than that used for administration (or a composition thereof with other DNA fragments) may be used as an exogenous DNA sequence, and it is intended that the region (or composition of regions) be incorporated at the site of a double-stranded break in the target DNA induced by the PpCas9 nuclease. In some embodiments of the present invention, a region of double-stranded DNA of the genomic DNA of an organism used for the introduction of the PpCas9 protein may be further modified by mutation (nucleotide substitution) and insertion or deletion of one or more nucleotides and used as an exogenous DNA sequence.

[0017] The technical result of the present invention is to increase the versatility of the available CRISPR-Cas9 system and enable the use of the Cas9 nuclease to cleave genomic or plasmid DNA at more specific sites and under more specific conditions. The novel nuclease can be used in cells of bacteria, mammals, or other organisms.

Brief Description of the Drawings

[0018]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Modes for Carrying Out the Invention

[0019] The terms "includes" and "including" used in the description of the present invention shall be construed to mean "including, among other things". The said terms are not intended to be construed as "consisting only of". Unless otherwise defined, technical and scientific terms within this application have the typical meanings generally accepted in scientific and technical literature.

[0020] As used herein, the term "percent homology of two sequences" is equal to the term "percent identity of two sequences". Sequence identity is determined based on a reference sequence. Algorithms for sequence analysis are known to those of skill in the art, such as BLAST described in Altschul et al., J. Mol. Biol. 215, 403-410 (1990). For the purposes of determining the level of identity and similarity between nucleotide sequences as well as between amino acid sequences in the present invention, comparison of nucleotide and amino acid sequences may be used, and the comparison is performed using gap alignment with standard parameters by the BLAST software package provided by the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / blast). The percent identity of two sequences is determined by the number of positions of identical amino acids in these two sequences, taking into account the number of gaps inserted for optimal comparison of the two sequences by alignment and the length of each gap. Percent identity is the same as dividing the number of identical amino acids at a given position by the total number of positions, considering the sequence alignment, and multiplying by 100.

[0021] The term "specifically hybridizes" refers to the association between two single-stranded nucleic acid molecules or between sequences that are sufficiently complementary to permit such hybridization under established conditions commonly used by those of skill in the art.

[0022] The phrase "double-strand break located immediately prior to the nucleotide PAM sequence" means that the double-strand break within the target DNA sequence will be made 0 to 25 nucleotides away from the nucleotide PAM sequence towards the 5' end.

[0023] The exogenous DNA sequence introduced simultaneously with the guide RNA is intended to refer to a DNA sequence specifically prepared to specifically modify the double-stranded target DNA at the cleavage site determined by the specificity of the guide RNA. Such modification may be, for example, an insertion or deletion of a specific nucleotide at the cleavage site within the target DNA. The exogenous DNA may be either a DNA region from a different organism or a DNA region from the same organism as the target DNA.

[0024] A protein containing a specific amino acid sequence shall refer to a protein having an amino acid sequence composed of the amino acid sequence and, optionally, other sequences linked to the amino acid sequence by peptide bonds. Examples of other sequences may be a nuclear localization signal (NLS) or other sequences that bring about an increase in the function of the amino acid sequence.

[0025] The exogenous DNA sequence introduced simultaneously with the guide RNA shall refer to a DNA sequence specifically prepared to specifically modify double-stranded target DNA at a cleavage site determined by the specificity of the guide RNA. Such a modification may be, for example, the insertion or deletion of specific nucleotides at the cleavage site within the target DNA. The exogenous DNA may be either a DNA region derived from a different organism or a DNA region derived from the same organism as the target DNA.

[0026] The effective amounts of the protein and RNA introduced into the cell shall refer to the amounts of the protein and RNA that, when introduced into the cell, can form a functional complex, that is, a complex that specifically binds to the target DNA and causes a double-strand break at a site determined by the guide RNA and the PAM sequence on the DNA. The efficiency of this process can be evaluated by analyzing the target DNA isolated from the cell using conventional techniques known to those skilled in the art.

[0027] The protein and RNA can be delivered into the cell by various techniques. For example, the protein may be delivered as a DNA plasmid encoding the gene of this protein, an mRNA translated into this protein in the cytoplasm, or a ribonucleoprotein complex containing this protein and the guide RNA. The delivery can be carried out by various techniques known to those skilled in the art.

[0028] Nucleic acids encoding components of the system can be introduced into cells directly or indirectly by, for example, transfection or transformation of cells by methods known to those skilled in the art, use of recombinant viruses, manipulation of cells such as DNA microinjection, and the like.

[0029] A ribonucleic acid complex consisting of a nuclease, guide RNA, and exogenous DNA (if necessary) may be delivered by transfecting the complex into cells or mechanically introducing the complex into cells, for example, by microinjection.

[0030] The nucleic acid molecule encoding the protein to be introduced into the cell may be integrated into the chromosome or may be DNA that replicates extrachromosomally. In some embodiments, to ensure efficient expression of the protein gene by the DNA introduced into the cell, it is necessary to modify the sequence of the DNA to optimize the codons for expression, depending on the cell type, because the frequency of synonymous codon occurrences in the coding regions of the genomes of various organisms is uneven. Codon optimization is necessary to increase expression in animal, plant, fungal, or microbial cells.

[0031] In the case of a protein having a sequence that is at least 95% identical to the amino acid sequence of SEQ ID NO: 1, in order to function in a eukaryotic cell, it is necessary for this protein to ultimately go into the nucleus of this cell. Thus, in some embodiments of the present invention, a protein having a sequence that is at least 95% identical to the amino acid sequence of SEQ ID NO: 1 and further modified at one or both ends by the addition of one or more nuclear localization signals is used to form a double-strand break in the target DNA. For example, a nuclear localization signal derived from the SV40 virus may be used. To provide efficient delivery to the nucleus, the nuclear localization signal may be separated from the main protein sequence by, for example, the spacer sequence described in Shen B, et al., "Generation of gene-modified mice via Cas9 / RNA-mediated gene targeting", Cell Res. May 2013; 23(5):720-723. Further, in other embodiments, different nuclear localization signals or alternative methods for delivering the protein to the cell nucleus may be used.

[0032] The present invention encompasses introducing double-strand breaks into DNA molecules at precisely specified positions using a protein derived from P. pneumotropica organisms that is homologous to the Cas9 protein characterized heretofore. Using CRISPR nucleases to introduce targeted modifications into the genome has many advantages. First, the specificity of the activity of the system is determined by the crRNA sequence, making it possible to use one type of nuclease for all target loci. Second, this technology enables the delivery of several guide RNAs complementary to different gene targets into the cell at once, thereby making it possible to modify several genes simultaneously at once.

[0033] PpCas9 is a Cas nuclease found in Pasteurella pneumotropica ATCC 35149, a pathogenic bacterium of rodents that inhabits the lungs of animals. The Pasteurella pneumotropica (P. pneumotropica) CRISPR Cas9 system (hereinafter referred to as CRISPR PpCas9) belongs to the type II-C CRISPR Cas system and consists of a CRISPR cassette that carries four tandem repeats (DRs) with the sequence 5’ATTATAGCACTGCGAAATGAAAAAGGGAGCTACAAC3’ separated by unique spacer sequences. None of the spacers of the system match the sequences of currently known bacteriophages or plasmids, and this fact makes it impossible to determine the PpCas9 PAM of interest by bioinformatics analysis. Adjacent to the CRISPR cassette are the gene for the effector Cas9 protein, PpCas9, and the genes for the Cas1 and Cas2 proteins involved in the adaptation and integration of new spacers. Near the Cas genes, a sequence was found that is partially complementary to the tandem repeats and folds into a characteristic secondary structure, and this sequence is thought to be the tracer RNA (tracrRNA) (Figure 1).

[0034] Findings on the characteristic composition of the RNA-Cas protein complex of the type II-C system made it possible to predict the direction of transcription of the CRISPR cassette: the pre-crRNA is transcribed in the direction opposite to the Cas genes (Figure 1).

[0035] Therefore, analysis of the sequence of the PpCas9 locus made it possible to predict the sequences of the tracer and guide RNAs (Table 1).

[0036]

Table 1

[0037] To verify the activity of PpCas9 nuclease and determine the PpCas9 PAM of interest, experiments were conducted to reproduce DNA cleavage reactions in vitro. To determine the PAM sequence of the PpCas9 protein, in vitro cleavage of a double-stranded PAM library was utilized. For this purpose, it was necessary to obtain all components of the PpCas9 effector complex: guide RNA and nuclease, in recombinant form. The determination of the guide RNA sequence enabled the synthesis of crRNA and tracrRNA molecules in vitro. The synthesis was carried out using the NEB HiScribe T7 RNA Synthesis Kit. The double-stranded DNA library was a 374-base pair (bp) fragment containing a protospacer sequence flanked by a random 7-nucleotide (5’-NNNNNNN-3’) from the 3’ end side:

[0038] [Chemical Formula]

[0039] It was as follows. To cleave this target, guide RNAs with the following sequences: tracrRNA:

[0040] [Chemical Formula]

[0041] and crRNA:

[0042] [Chemical Formula]

[0043] were used. The bold font indicates the crRNA sequence complementary to the protospacer (target DNA sequence). To produce the recombinant PpCas9 protein, its gene was cloned into the plasmid pET21a. DNA synthesized by Integrated DNA Technologies (IDT) was used as the DNA encoding the gene. The sequence was codon-optimized to exclude rare codons found in the P.neumotropica genome. Escherichia coli (E.coli) Rosetta cells were transformed with the obtained plasmid pET21a-6×His-PpCas9.

[0044] 500 μL of the overnight culture was diluted into 500 mL of LB medium, and the cells were grown at 37 °C until an optical density OD600 of 0.6 was obtained. The synthesis of the target protein was induced by adding IPTG at a concentration of 1 mM, and then the cells were incubated at 20 °C for 6 hours. The cells were then centrifuged at 5000 g for 30 minutes, and the obtained cell pellet was frozen at -20 °C.

[0045] The pellet was thawed on ice for 30 minutes and resuspended in 15 mL of lysis buffer (50 mM Tris-HCl, pH 8, 500 mM NaCl, 1 mM β-mercaptoethanol, 10 mM imidazole) supplemented with 15 mg of lysozyme, and reincubated on ice for 30 minutes. The cells were then disrupted by sonication for 30 minutes and centrifuged at 16000 g for 40 minutes. The obtained supernatant was passed through a 0.2 μm filter and applied to a HisTrap HP 1 mL column (GE Healthcare) at a rate of 1 mL / min.

[0046] Chromatography was performed at a rate of 1 mL / min using an AKTA FPLC chromatograph (GE Healthcare). The column containing the applied protein was washed with 20 mL of lysis buffer supplemented with 30 mM imidazole, and then the protein was eluted with lysis buffer supplemented with 300 mM imidazole.

[0047] Next, the protein fraction obtained in the process of affinity chromatography was passed through a Superdex 200 10 / 300 GL gel filtration column (24 mL) equilibrated with buffer: 50 mM Tris-HCl pH8, 500 mM NaCl, 1 mM DTT. Using an Amicon concentrator (including a 30 kDa filter), the fraction corresponding to the monomeric form of the PpCas9 protein was concentrated to 3 mg / mL, and then the purified protein was stored at -80 °C in a buffer containing 10% glycerol.

[0048] The in vitro reaction to cleave the linear PAM library was carried out in a volume of 20 μL under the following conditions. The reaction mixture consisted of 1× CutSmart buffer (NEB), 5 mM DTT, 100 nM PAM library, 2 μM trRNA / crRNA, and 400 nM PpCas9 protein. As a control, a sample without RNA was prepared in a similar manner. The samples were incubated at different temperatures and analyzed by gel electrophoresis in a 2% agarose gel. When the DNA is correctly recognized and specifically cleaved by the PpCas9 protein, two DNA fragments of approximately 326 and 48 base pairs should be generated (see Figure 2).

[0049] The experimental results showed that PpCas9 has nuclease activity and cleaves a part of the PAM library fragment. The temperature gradient showed that the protein is active in the temperature range of 35 - 45 °C (Figure 3). This study then used a temperature of 42 °C as the working temperature.

[0050] The library cleavage reaction was repeated under selected conditions. The reaction products were applied to a 1.5% agarose gel and subjected to electrophoresis. The intact DNA fragment of 374 bp in length was extracted from the gel and prepared for high-throughput sequencing using the NEB NextUltra II kit. The samples were sequenced on an Illumina platform, and then sequence analysis was performed using bioformatical methods: (Maxwell CS, et al., A detailed cell-free transcription-translation-based assay to decipher CRISPR protospacer-adjacent motif. Methods. July 1, 2018; 143:48-57) to determine the difference in the nucleotide occurrence rate at each position of the PAM (NNNNNNN) compared to the control samples. Furthermore, a PAM logo was created to analyze the results (Figure 4).

[0051] Both methods of data analysis showed significance at positions 5, 6, and 7 of the PAM (Figure 4). Therefore, through in vitro analysis, the putative PAM sequence for PpCas9 could be established as NNNNATT. However, this sequence is only an estimate considering the inaccurate results obtained by the screening method for determining the PAM.

[0052] In this regard, the significance of the positions of individual PAM sequences was verified to more accurately determine the sequence. For this purpose, a DNA fragment containing the DNA target 5’-atctcctttcattgagcac-3’ adjacent to the PAM sequence

[0053]

Chemical formula

[0054] (or its derivative):

[0055]

Chemical formula

[0056] performed an in vitro reaction of cleavage. All DNA cleavage reactions were performed under the following conditions: 1× CutSmart buffer 400 nM PpCas9 20 nM DNA 2 μM crRNA 2 μM tracrRNA Incubation time 30 minutes, reaction temperature 42 °C.

[0057] Substitutions at the PAM position 1 by all four possible nucleotide variants did not affect the efficiency of protein activity (Figure 5). The predicted significance at positions 5 and 6 was experimentally confirmed by single nucleotide substitutions (purine by pyrimidine and vice versa) at each of the PAM positions. When the substitution occurred at positions 5 and 6, the protein substantially stopped its activity. When the substitution occurred at position 7, the efficiency of PpCas9 activity decreased by half, a fact that reflects a reduced requirement for the nucleotide at this position (Figure 6). Therefore, according to the results of in vitro PAM screening of PpCas9 nuclease, the most likely nucleotides at the PAM position 5 are adenine or guanine, a fact that was experimentally confirmed (Figure 7). The substitution from A to G did not reduce the efficiency of fragment cleavage.

[0058] According to the results of in vitro screening, fragments with "T" or "S" at the 7th position should be recognized more efficiently. Further experiments were conducted to finally verify the significance of the nucleotides at this position. The results of the in vitro test showed that substitution of the nucleotide "T" at the 7th position with A or G reduced the cleavage efficiency by 40 - 50% (Figure 8). Therefore, the 7th position of PAM is not well conserved compared to the 5th and 6th positions: purines at the 7th position reduce the recognition efficiency but do not prevent the PpCas9 protein from introducing double-strand breaks into DNA.

[0059] The results of the study were as follows: The PAM recognized by the PpCas9 nuclease corresponds to the formula 5'-NNNN(A / G)TT-3'. The 7th position is not well conserved. The present invention includes the following aspects non-limitingly. [Aspect 1] Use of a protein comprising the amino acid sequence of SEQ ID NO: 1, or an amino acid sequence that is at least 95% identical to the amino acid sequence of SEQ ID NO: 1 and differs from SEQ ID NO: 1 only at non-conserved amino acid residues, for forming a double-strand break in a DNA molecule immediately preceding the nucleotide sequence 5'-NNNN(A / G)TT-3' within the DNA molecule. [Aspect 2] Use of the protein according to Aspect 1, wherein the double-strand break in the DNA molecule is formed at a temperature of 35°C to 45°C. [Aspect 3] Use of the protein according to Aspect 1, wherein the protein comprises the amino acid sequence of SEQ ID NO: 1. [Aspect 4] Use of the protein according to Aspect 1, wherein the double-strand break in the DNA molecule is formed within the genomic DNA of mammalian cells. [Aspect 5] Use of the protein according to embodiment 4, characterized in that the double-strand break in the DNA molecule results in modification of the genomic DNA of the mammalian cell, the use of the protein. [Embodiment 6] A method for modifying the genomic DNA sequence of a cell of a unicellular or multicellular organism containing genomic DNA, comprising an effective amount of: a) a protein containing the amino acid sequence of SEQ ID NO: 1 or a nucleic acid encoding a protein containing the amino acid sequence of SEQ ID NO: 1, and b) a nucleotide sequence 5'-NNNN(A / G)TT-3' adjacent to the genomic DNA region of the organism forming a double strand with the nucleotide sequence and containing a sequence that interacts with the protein after double strand formation, or introducing the DNA sequence encoding the guide RNA into the cell of the organism, A method, wherein the interaction of the guide RNA and the nucleotide sequence 5'-NNNN(A / G)TT-3' with the protein results in the formation of a double-strand break in the genomic DNA sequence immediately adjacent to the sequence 5'-NNNN(A / G)TT-3'. [Embodiment 7] The method according to embodiment 6, further comprising the step of introducing an exogenous DNA sequence simultaneously with the guide RNA, the method. [Embodiment 8] The method according to embodiment 6, characterized in that the cell is a mammalian cell, the method.

[0060] The following exemplary embodiments of the method are provided for the purpose of disclosing the features of the present invention and should in no way be construed as limiting the scope of the present invention. Example 1 Testing the activity of the PpCas9 protein in the cleavage of various DNA targets.

[0061] To investigate the ability of PpCas9 to recognize various DNA sequences adjacent to the sequence 5'-NNNN(A / G)TT-3', experiments were performed on the in vitro cleavage of DNA targets derived from the human grin2b gene sequence (see Table 2).

[0062] [Table 2]

[0063] A PCR fragment of the grin2b gene carrying a recognition site (Table 2) that is probably recognizable by PpCas9 by the PAM consensus sequence 5'-NNNN(A / G)TT-3' was used as a target in the cleavage reaction. CrRNAs targeting PpCas9 to these sites were synthesized to recognize these sequences.

[0064] The cleavage reaction was carried out under the conditions selected for PpCas9; the results are shown in Figure 9. Figure 9 shows that the PpCas9 enzyme successfully cleaved three of the four targets with the appropriate PAM.

[0065] The target in lane 6 has the PAM sequence CAGCATT, and according to predictions based on the results of the removal analysis, this sequence should be efficiently recognized by the protein. However, recognition of this fragment did not occur in this experiment.

[0066] Therefore, the PAM CAGCATT was further verified against another protospacer target restricted by the same PAM (Figure 10). In this case, the PAM was effectively recognized, resulting in DNA cleavage. Therefore, the protein has some additional preference for DNA target sequences. The preference is probably related to the secondary structure of the DNA.

[0067] Therefore, this study demonstrated the presence of nuclease activity in PpCas9 and further enabled the determination of its PAM sequence and the verification of the guide RNA sequence. The PpCas9 ribonucleoprotein complex specifically introduces cleavage into targets restricted from the 5'-end of the protospacer to the PAM 5'-NNNNTT(A / G)-3'. An overview of the PpCas9 / RNA complex is shown in Figure 11.

[0068] Example 2 Use of hybrid guide RNAs to cleave DNA targets. sgRNA is a form of guide RNA that fuses tracrRNA (trans-activating crRNA) and crRNA. To select the optimal sgRNA, three variants of this sequence were constructed, which differed in the length of the tracrRNA-crRNA duplex. The RNA was synthesized in vitro and experiments containing it were performed for DNA target cleavage (Figure 12).

[0069] The following RNA sequences: 1-sgRNA1 25DR:

[0070]

Chem.

[0071] 2-sgRNA2 36DR

[0072]

Chem.

[0073] were used as hybrid RNAs. The bold font indicates the 20-nucleotide sequence (variable part of sgRNA) that provides pairing with the DNA target. Additionally, the experiments used RNA-free control samples and a positive control that cleaves the target using crRNA + trRNA.

[0074] Sequences containing the recognition site 5’tatctcctttcattgagcac3’ along with the corresponding common sequence PAM CAACATT:

[0075]

Chem.

[0076] were used as DNA targets. The bold font indicates the recognition site and the capital letters represent the PAM. The reaction was carried out under the following conditions: the concentration of the DNA sequence containing PAM (CAACATT) was 20 nM, the protein concentration was 400 nM, the RNA concentration was 2 μM; the incubation time was 30 minutes, and the incubation temperature was 37 °C.

[0077] The selected sgRNA1 and sgRNA2 were found to be as effective as the native tracrRNA and crRNA sequences: cleavage occurred in more than 80% of the DNA targets (Figure 12).

[0078] These hybrid RNA variants may be used to cleave any other target DNA after modifying the sequence that directly pairs with the DNA target. Example 3 A Cas9 protein derived from a related organism belonging to P.neumotropica.

[0079] To date, the CRISPR-Cas9 enzyme has not been characterized in P.neumotropica. The Cas9 protein from Staphylococcus aureus of a similar size is 28% identical to PpCas9 (Figure 13, the degree of identity was calculated with the default parameters of the BLASTp software). A similar degree of identity exists in other known Cas9 proteins (not shown).

[0080] Therefore, the PpCas9 protein is significantly different from other Cas9 proteins studied to date in its amino acid sequence. One of ordinary skill in the art of genetic engineering will recognize that the PpCas9 protein sequence variants obtained and characterized by the applicant in this description can be modified without changing the function of the protein itself [e.g., by site-directed mutagenesis of amino acid residues that do not directly affect functional activity (Sambrook et al., Molecular Cloning: A Laboratory Manual, (1989), CSH Press, pp. 15.3-15.108)]. In particular, one of ordinary skill in the art will recognize that non-conserved amino acid residues can be modified without affecting residues that are involved in (determine) protein function or structure. Examples of such modifications include substitution of non-conserved amino acid residues with homologous amino acids. A portion of the region containing non-conserved amino acid residues is shown in Figure 12. In some embodiments of the invention, a protein comprising an amino acid sequence that is at least 95% identical to the amino acid sequence of SEQ ID NO: 1 and that differs from SEQ ID NO: 1 only at non-conserved amino acid residues is used to form a double-strand break within a DNA molecule immediately prior to the nucleotide sequence 5'-NNNN(A / G)TT-3' within the DNA molecule. Homologous proteins can be obtained by mutagenesis of the corresponding nucleic acid molecule (e.g., site-specific or PCR-mediated mutagenesis), and then the encoded modified Cas9 protein can be tested for retention of its function according to the functional assays described herein.

[0081] Example 4 Modification of genomic DNA of human cells using PpCas9. To modify the genomic DNA of human cells, the PpCas9 nuclease gene was cloned under the control of the CMV promoter in a eukaryotic plasmid vector. Sequences encoding nuclear localization signals that ensure delivery of the nuclease to the cell nucleus were added to the 5' and 3' ends of the PpCas9 gene. The sgRNA sequence was cloned into the vector under the control of the U6 promoter. To test the activity of the system, sgRNAs with sequences complementary to target DNAs 20 and 24 nucleotides in length were used. A similar plasmid carrying a genomic DNA modification system based on SpCas9 known from the state of the art was used as a positive control. To evaluate the effectiveness of transfection, the plasmid further carried the GFP (green fluorescent protein) gene. The following regions of the human genomic DNA were used as DNA targets (Table 3).

[0082]

Table 3

[0083] For effective activity of the PpCas9 nuclease in eukaryotic cells, it is necessary to import the protein into the nucleus of the eukaryotic cells. This may be done by using an SV40 T antigen-derived nuclear localization signal (Lanford et al., Cell, 1986, 46: 575-582) linked to the PpCas9 sequence either via a spacer sequence or without a spacer sequence as described in Shen B, et al. "Generation of gene-modified mice via Cas9 / RNA-mediated gene targeting", Cell Res. May 2013; 23(5): 720-723.

[0084] In a given example, the complete amino acid sequence of the nuclease transported into the nucleus of human cells is the following sequence:

[0085]

Chemical formula

[0086] It was The plasmid used in this experiment has the sequence:

[0087]

Chemical formula

[0088]

Chemical formula

[0089] It had The following parts were distinguished within the plasmid sequence: U6 promoter (first region, uppercase), sequence complementary to the protospacer ("XXX-XXX"), conserved part of the sgRNA (third region, uppercase), PpCas9 gene (bold and emphasized), GFP gene (last region, uppercase).

[0090] Plasmids carrying PpCas9 or SpCas9 were transfected into human HEK293T cell cultures using Lipofectamine 2000 reagent. Cells were lysed 72 hours after transfection, and the resulting lysates were subjected to PCR to generate regions containing the target modification sites of genomic DNA. The resulting PCR fragments were subjected to an in vitro reaction with T7 endonuclease I to determine the frequency of insertions and deletions at the target sites of genomic DNA. The reaction products were applied to an agarose gel and subjected to electrophoresis. Figure 14A shows that PpCas9 can actively introduce modifications into the EMX1 and GRIN2b genes with an efficiency similar to that of the SpCas9 nuclease described in the prior art.

[0091] This experiment showed that PpCas9 requires an extended sgRNA compared to SpCas9 to effectively modify genomic DNA: in a given example, when using an sgRNA containing a sequence complementary to a DNA target with a length of 24 nucleotides, the efficiency of genetic modification is greater (compared to a length of 20 nucleotides).

[0092] High-throughput array determination was used to confirm the modifications introduced into the target DNA site. Figure 14B shows an example of a detectable modification of the nucleotide sequence of the EMX1 gene. Delivery of NLS_PpCas9_NLS to human cells may be utilized by taking advantage of delivery in the form of ribonucleic acid complexes. Delivery is carried out by incubating the recombinant form of PpCas9 NLS with guide RNA in CutSmart buffer (NEB). The recombinant protein is prepared from bacterial production cells by purification by affinity chromatography (NiNTA, Qiagen) and size exclusion chromatography (Superdex 200).

[0093] The protein is mixed with RNA at a ratio of 1:2 (PpCas9 NLS:sgRNA), the mixture is incubated at room temperature for 10 minutes, and then transfected into cells. Next, the DNA extracted therefrom is analyzed for insertions / deletions at the target DNA site (as described above).

[0094] The PpCas9 nuclease derived from the bacterium Pasteurella pneumotropica characterized in the present invention can be delivered to cells of various origins using standard techniques and methods known to those skilled in the art for modifying DNA. PpCas9 has many advantages compared to Cas9 proteins characterized heretofore.

[0095] Unlike other known Cas nucleases, PpCas9 has a short two-letter PAM required for the system to function. The present invention has shown that the presence of a short PAM (RTT) located 4 nucleotides away from the protospacer is sufficient for PpCas9 to function successfully in vivo.

[0096] Many of the previously known small-sized Cas nucleases that have the ability to introduce double-strand breaks into DNA have complex multi-character PAM sequences, limiting the options for sequences suitable for cleavage. Among the Cas nucleases studied to date that recognize short PAMs, only PpCas9 can recognize sequences adjacent to the RTT motif.

[0097] A second advantage of PpCas9 is its small protein size (1055 a.a.r). To date, it is the only small-sized protein studied that has a three-character RTT PAM sequence.

[0098] PpCas9 is a novel, small-sized Cas nuclease with a short and user-friendly PAM that is different from the PAM sequences of other currently known nucleases. The PpCas9 protein can efficiently cleave various DNA targets, including genomic DNA of human cells, at 37°C, and can serve as the basis for a new genome editing tool.

[0099] The present invention has been described with reference to the disclosed embodiments, but those skilled in the art will recognize that the specific embodiments described in detail are provided for the purpose of exemplifying the present invention and should in no way be construed as limiting the scope of the present invention. It will be understood that various modifications can be made without departing from the spirit of the present invention.

Claims

1. Use of a protein in vitro for forming a double-strand break at a position immediately before the nucleotide sequence 5'-NNNN(A / G)TT-3' within a DNA molecule, wherein said protein comprises the amino acid sequence of SEQ ID NO:

1.

2. The use according to claim 1, wherein the double-strand break within the DNA molecule is formed at a temperature of 35°C to 45°C.

3. The use according to claim 1, wherein the double-strand break within the DNA molecule is formed within the genomic DNA of mammalian cells.

4. The use according to claim 3, wherein the double-strand break within the DNA molecule results in modification of the genomic DNA of the mammalian

5. A method for modifying the genomic DNA sequence of a cell of a unicellular or multicellular organism containing genomic DNA in vitro, comprising introducing into the cell of the organism an effective amount of: a) a protein comprising the amino acid sequence of SEQ ID NO: 1 or a nucleic acid encoding a protein comprising the amino acid sequence of SEQ ID NO: 1, and b) a guide RNA comprising a sequence that forms a double strand with the nucleotide sequence of a genomic DNA region of the organism at a position immediately before the nucleotide sequence 5'-NNNN(A / G)TT-3' and interacts with said protein after double-strand formation, or DNA encoding said guide RNA, wherein the interaction of the guide RNA and the nucleotide sequence 5'-NNNN(A / G)TT-3' with the protein results in the formation of a double-strand break in the genomic DNA sequence at a position immediately before the sequence 5'-NNNN(A / G)TT-3'.

6. The method according to claim 5, further comprising introducing exogenous DNA simultaneously with the guide RNA.

7. The method according to claim 5, wherein the cell is a mammalian cell. The method.

Citation Information

Patent Citations

  • Novel CAS9 orthologs

    WO2019165168A1