DNA methyltransferase
Patent Information
- Application Number
- JP2022087994
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-05-30
- Publication Date
- 2025-06-09
AI Technical Summary
Current DNA methyltransferases (MTases) used for epigenetic mapping in mammalian cells have recognition sequences that overlap with the CG dinucleotide, limiting their ability to provide fine resolution and distinguishable methylation patterns.
Development of a novel DNA methyltransferase that specifically recognizes and methylates the CC dinucleotide sequence, avoiding overlap with CG dinucleotides, by modifying the N-terminus of existing MTases such as M.CviQIX and M.CviPII to enhance methylation activity.
Enables simultaneous profiling of endogenous methylation and other epigenomes with higher resolution, allowing for distinct methylation patterns to be detected without fragmentation, facilitating advanced epigenome analysis.
Smart Images

Figure 00000021_0000 
Figure 00000021_0001 
Figure 00000022_0000
Abstract
Description
[Technical Field]
[0001] This invention relates to a novel DNA methyltransferase. [Background technology]
[0002] DNA methyltransferases (MTases) are enzymes that induce methylation, one of the representative epigenetic modifications of DNA, and are useful tools for epigenomic analysis. DNA methylation at the 5-position of cytosine (DNA cytosine 5-methylation (5mC)) is widely observed in the genomic DNA of living organisms. However, with respect to DNA cytosine 5-methylation, the MTases identified so far (Non-Patent Documents 1-4) have drawbacks, such as their recognition sequences overlapping with the CG dinucleotide sequence, which is the recognition sequence for intrinsic DNA methylation (endogenous methylation) in mammalian cells, and the fact that the number of bases in the recognition sequence is too long for fine epigenetic mapping. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Renbaum, P. et al., Cloning, characterization, and expression in Escherichia coli of the gene coding for the CpG DNA methylase from Spiroplasma sp. strain MQ1(M.SssI). Nucleic acids research, 1990. 18(5): p. 1145-1152. [Non-Patent Document 2] Xu, M. et al., Cloning, characterization and expression of the gene coding for a cytosine-5-DNA methyltransferase recognizing GpC. Nucleic Acids Res, 1998. 26(17): p. 3961-3966. (Same as reference 21 below) [Non-Patent Document 3] Zhang, B. et al., The M.AluI DNA-(cytosine C5)-methyltransferase has an unusually large, partially dispensable, variable region. Nucleic Acids Res, 1993. 21(4): p. 905-911. [Non-Patent Document 4] Vladimir S Dedkov et al., Cloning and Study of New DNA Methyltransferase M.FatI Modifying Cytosine in a Recognition Site CATG. Research Journal of Pharmaceutical, Biological and Chemical Sciences, 2015. 6(6): p. 1341-1348. [Overview of the Initiative] [Problems that the invention aims to solve]
[0004] Under these circumstances, there was a need for the development of DNA methyltransferases (MTases) that could introduce methylation distinct from endogenous methylation, that is, that could recognize and methylate short contexts of similar length to CG dinucleotide sequences other than those where endogenous methylation occurs. [Means for solving the problem]
[0005] This invention has been made in consideration of the above circumstances and provides the following: a protein (MTase) and a gene encoding it.
[0006] [1] The following proteins: (a), (b), or (c). (a) A protein obtained by deleting some amino acid residues, including the N-terminus, of the DNA methyltransferase M.CviQIX or M.CviPII, thereby acquiring the activity to specifically recognize CC dinucleotide sequences and 5-methylate the cytosine residue at the 5' end of said sequence. (b) A protein having an amino acid sequence in which one or more amino acids are deleted, substituted, or added in the amino acid sequence of the protein described in (a) above, and which has the activity to specifically recognize a CC dinucleotide sequence and 5-methylate the cytosine residue at the 5' end of the sequence. (c) A protein having an amino acid sequence that is 80% or more identical to the amino acid sequence of the protein in (a) above, and that has the activity to specifically recognize the CC dinucleotide sequence and 5-methylate the cytosine residue at the 5' end of the sequence.
[0007] [2] The following proteins: (a), (b), or (c). (a) A protein having the amino acid sequence shown in SEQ ID NO: 6 or 8. (b) A protein having an amino acid sequence in which one or more amino acids are deleted, substituted, or added in the amino acid sequence of the protein described in (a) above, and which has the activity to specifically recognize a CC dinucleotide sequence and 5-methylate the cytosine residue at the 5' end of the sequence. (c) A protein having an amino acid sequence that is 80% or more identical to the amino acid sequence of the protein in (a) above, and that has the activity to specifically recognize the CC dinucleotide sequence and 5-methylate the cytosine residue at the 5' end of the sequence.
[0008] [3] A gene encoding the protein described in [1] or [2] above. [4] A gene containing DNA of (a), (b), or (c) below. (a) DNA containing the nucleotide sequence shown in Sequence ID No. 5 or 7. (b) DNA that hybridizes under stringent conditions with DNA having a base sequence complementary to the DNA in (a) above, and that encodes a protein that specifically recognizes a CC dinucleotide sequence and has the activity to 5-methylate the cytosine residue at the 5' end of the sequence. (c) DNA that has 80% or more identity with the DNA in (a) above, and that encodes a protein that specifically recognizes the CC dinucleotide sequence and has the activity to 5-methylate the cytosine residue at the 5' end of the sequence.
[0009] [5] Recombinant vectors containing the gene described in [3] or [4] above. [6] A transformant comprising the recombinant vector described in [5] above. [7] A step of culturing the transformant described in [6] above, The process involves collecting a protein from the resulting culture that has the activity to specifically recognize the CC dinucleotide sequence and 5-methylate the cytosine residue at the 5' end of the sequence. A method for producing the protein, including the protein. [8] An epigenome analysis method comprising using the protein described in [1] or [2] above, or a protein produced by the method described in [7] above. [Effects of the Invention]
[0010] According to the present invention, it is possible to introduce methylation that is distinguishable from intrinsic DNA methylation (endogenous methylation) in mammalian cells, and more specifically, a protein (MTase) and a gene encoding it that can specifically recognize a "CC dinucleotide sequence" and 5-methylate the cytosine residue on the 5' side of the sequence. The protein (MTase) of the present invention is useful in that it theoretically enables epigenome analysis methods (epigenome mapping) that simultaneously profile endogenous methylation and other epigenome processes. [Brief explanation of the drawing]
[0011] [Figure 1] It is a diagram showing the scheme of plasmid vector pCZS. [Figure 2] It is a diagram showing the dual affinity purification strategy used to purify candidate MTases. A. Structure of the expressed protein. B. Proteins purified by HisTrap purification and dual affinity purification combined with HisTrap and StrepTrap purification. Here, M.CviPI is shown as an example of MTase. C. The DNA methylation activity of purified M.CviPI was assayed using the methylation-sensitive restriction enzyme HaeIII. Unmethylated λDNA (Promega) was used as the substrate DNA. As a control, a lane using M.CviPI purchased from NEB was placed. [Figure 3] It is a diagram showing that the N-terminal extension of M.CviQIX has an inhibitory effect on DNA methylation. A. Diagram showing the recombinant protein structure of M.CviQIX. B. Recognition sequence of the restriction enzyme HaeIII. The recognition sequence contains the CC dinucleotide. If the first C is 5-methylated, HaeIII cannot cleave the substrate DNA. C. Diagram showing the results of digesting genomic DNA extracted from bacterial cells expressing various M.CviQIX proteins shown in A with HaeIII and analyzing it on an E-Gel Ex. D. Scheme showing the structures of various M.CviQIX protein constructs. The tag means a sequence derived from the vector. E. Genomic DNA digestion of bacterial cells using the construct of D. [Figure 4] It is a diagram showing the experimental strategy for clarifying the recognition sequence of MTase candidates. A. The scheme used. B. Examples of the identified sequence motifs. C. A phylogenetic tree was drawn based on the similarity of the amino acid sequences. The determined recognition sequences are arranged according to the gene names. The common recognition sequences of each clade are annotated at the top of the clade. [Figure 5]This figure shows the results of NOMe-Seq on fixed yeast nuclei. A. Comparison of methylation levels between CCMT-treated and GCMT-treated nuclei. The bin size is 10 bp. B. Yeast gene (Track 1), DNA methylation levels at GC sites (Track 2, NOMe-Seq using GCMT) and CC sites (Track 3, NOMe-Seq using CCMT), ATAC-Seq read coverage (Track 4, SRR12926697) (Reference 25), and MNase-Seq read coverage (Track 5, SRR6729489) (Reference 26). CE. Aggregation plot around the start codon of the yeast gene. Examples of plotting DNA methylation signals at CC (C) and GC (D) sites, as well as MNase-Seq (SRR6729489) are shown. [Figure 6] This figure shows the results of NOMe-Seq of IMR-90 cells. AD. Comparison of DNA methylation levels between CCMT-treated and untreated nuclei. Endogenous DNA methylation at CG sites (A and B) and methylation at CC sites (C and D) introduced in vitro via CCMT are shown. E. A genome braser screenshot (Track 8) shows human genes (Track 1), intrinsic methylation levels (untreated, CCMT-treated, and GCMT-treated, respectively), MTase-induced methylation (Untreated, CCMT-treated, and GCMT-treated, respectively), and ATAC-Seq read coverage (SRR7765313). The figure on the right is a magnified view of the indicated location in the figure on the left. For tracks showing MTase-induced methylation patterns (Tracks 5, 6, and 7), the peaks detected by macs2 are shown at the bottom of each track. [Figure 7]This is a figure showing attempts to express M.CviQIX and M.CviPII from several expression vectors. A. Schematic of the construct used for the expression of M.CviQIX. B. Recognition sequences of restriction enzymes used to examine the methylation state of genomic DNA. When the underlined C is methylated, the activity of the enzyme is inhibited. C. Analysis of the methylation state of genomic DNA extracted from Escherichia coli cells in which the expression of M.CviQIX shown in A was induced. D. Schematic of the N-terminal deletion mutant of M.CviQIX. A construct based on the cold shock promoter was used. E. Restriction digestion assay of genomic DNA extracted from Escherichia coli cells having the construct shown in D. F. Deletion mutants of M.CviPII were compared. G. Restriction digestion assay of genomic DNA extracted from Escherichia coli cells containing the construct of F. [Figure 8] This is a figure showing the purification of MTase and its activity. A. Representative SDS-PAGE gel image analyzing protein purification. W: Whole cell extract; 3 and 4: The third and fourth fractions of HisTrap column purification, respectively. I: Input of StrepTrap purification. S: Purified protein. B. Analysis of the MTase activity of the purified protein. Unmethylated λDNA was treated with the purified MTase, digested with the methylation-sensitive restriction enzyme HpaII, and analyzed by E-Gel Ex. [Figure 9] This is a figure showing the results of examining the optimal salt concentration of CCMT. A. Strategy used for the analysis. B. Methylation levels observed under the salt conditions used. [Figure 10] This is a figure showing the analysis results by GCMT- and CCMT-based NOMe-Seq of fixed budding yeast nuclei. A. Aggregation plots regarding methylation levels of GCMT-, CCMT-based NOMe-Seq, and MNaseseq. B-C. Venn diagrams showing overlapping peaks identified by GCMT-based, CCMT-based NOMe-Seq, ATAC-seq, and MNase-seq. **Modes for Carrying Out the Invention**
[0012] The present invention will now be described in detail. The scope of the present invention is not limited to this description, and modifications can be made to the extent that the spirit of the invention is not impaired, in addition to the examples given below. All publications cited herein, such as prior art documents, and published gazettes, patent gazettes, and other patent documents, are incorporated herein by reference. In this specification, "DNA methyltransferase," "DNA methyltransferase," and "MTase" are synonymous. Further details of each reference cited in the following text are listed at the end of this specification.
[0013] 1. Overview and technical background of the present invention DNA cytosine 5-methylation (5mC) is widely observed in genomic DNA in living organisms and forms one of the major epigenetic modifications. Classically, in mammalian cells, increased 5mC levels were thought to lead to gene silencing, but it has been shown that methylation levels of the gene itself are positively correlated with gene expression levels (references 1-4). Therefore, the distribution pattern of 5mC is cell type specific, and the whole-genome distribution of 5mC or methylome has always attracted the interest of many researchers.
[0014] In mammalian cells, 5-methylation of cytosine depends on three types of DNA methyltransferases (MTases): DNMT1, DNMT3A, and DNMT3B (references 5-7). DNMT1 is a maintenance DNA methyltransferase that introduces methylation to hemimethylated CG dinucleotides. On the other hand, DNMT3A and DNMT3B are de novo methyltransferases that introduce DNA methylation to unmethylated CG sites. Although some exceptions are known in pluripotent stem cells, oocytes, neurons, and glial cells, DNA methylation to sites other than CG is rare in normal somatic tissues (references 8,9). Therefore, it is generally accepted that 5mC is introduced almost exclusively to CG dinucleotides, and cytosine in situations other than CG dinucleotides remains unmethylated in most mammalian cells.
[0015] The genomic placement of 5mC is closely related to other epigenetic modifications. For example, 5mC is depleted in gene promoters, and trimethylation of lysine 4 of histone H3 (H3K4me3) is enriched (Reference 10). Another example is the dimethylation and trimethylation of histone H3 lysine 36 (H3K36me2 and H3K36me3), whose genomic placement is positively related to transcriptional activity. H3K36me2 and H3K36me3 are gene-body enriched and introduce 5mC by recruiting DNMT3A / B. Therefore, these modifications support gene-body methylation (References 11,12). Since many epigenetic modifications are interrelated, detecting DNA methylation simultaneously with other epigenetic markers on the same DNA molecule is intriguing, and the realization of such measurements is expected to shift epigenetic measurement to the next stage. However, most current methods for detecting the epigenome are not suitable for simultaneous detection.
[0016] For example, consider chromatin immunoprecipitation followed by sequencing (ChIP-Seq) (references 13, 14), the most commonly used technique for detecting epigenomes. In ChIP-Seq, chromatin is first fragmented, then the target epigenome on the fragmented chromatin is immunoprecipitated, and the DNA bound to the epigenome is sequenced. Because ChIP-Seq is based on chromatin fragmentation, epigenomes located proximal to different nucleosomes should be separated in this initial fragmentation stage. Therefore, measuring the relationship between two epigenetic modifications on a single DNA molecule is not practical with ChIP-Seq. Consequently, fragmentation-based procedures are not suitable for simultaneously measuring multiple epigenetic modifications on a single DNA molecule.
[0017] However, there are several fragmentation-free procedures for epigenome detection. For example, DNA methyltransferase (MTase) can be used as a probe to measure chromatin accessibility. Nucleosome occupancy and methylome sequencing (NOMe-Seq) are among the earliest methods for measuring chromatin accessibility using MTase (Reference 15), and more recently, they have been performed in combination with nanopore sequencers (Reference 16). The same principle has recently been applied to Fiber-seq (Reference 17), which combines nonspecific adenine N-6 MTase M.EcoGII with a nanopore sequencer. Chromatin accessibility measures DNA-protein interactions of unspecified proteins, but interactions between DNA and specific proteins can also be detected with MTase. One representative technique is DamID (Reference 18). In DamID, dam methyltransferase fused with the target protein is expressed in target cells, and the specific interaction between the fusion protein and DNA can be measured by detecting DNA methylation introduced at the N-6 position of adenosine. In the initial DamID report (reference 18), methylation signals were detected using methylation-sensitive restriction enzymes, but this is not an absolute requirement for detecting DNA methylation signals. In fact, there have been cases where DNA methylation with DamID introduced has been detected using single-molecule sequencers. Furthermore, DiMeLo-Seq, which introduces adenine N-6 methylation near specific epigenetic modifications, has also been realized, further increasing the demand for (epigenetic) detection technologies using MTases.
[0018] Therefore, MTase enables simultaneous detection of multiple epigenomes on a single DNA molecule without fragmentation. In MTase-based epigenetic analysis, resolution is primarily determined by the frequency of recognized sequences in the genome, with the length of the recognized sequence determining its frequency. Since one of the most basic epigenetic units, the nucleosome, binds to approximately 150 base pairs (bp) of genomic DNA, the resolution required for epigenome mapping must be higher than 150 bp. Therefore, to obtain sufficient resolution for epigenetic measurements, the recognized sequence must be 3 bp (64 bp resolution) or less. Many MTases have been identified from bacteria and viruses, and non-specific adenine N-6 MTases such as M. EcoGII have achieved fine mapping of chromatin accessibility (reference 17). However, focusing on DNA cytosine C-5 MTases, the number of MTases available as probes for epigenome mapping is limited to only three: M. SssI ( C G), M. CviPI (G C ), M. CviPII ( C CD) (Target cytosine is underlined) (References 19, 20). Furthermore, since mammalian cells have intrinsic DNA methylation (endogenous methylation) of CG dinucleotides, the presence of sequences that completely (M.SssI) or partially (M.CviPI) overlap with CG dinucleotides is problematic. Therefore, it was desirable to develop and identify a cytosine C-5 MTase that does not overlap with CG dinucleotides and has a short recognition sequence.
[0019] The inventors have discovered a systematic strategy for identifying cytosine C-5 MTases in short recognition sequences, enabling the identification of MTases for six novel recognition sequences. Of these MTases, the inventors focused on those capable of recognizing CC dinucleotides, whose non-overlap specificity with CG dinucleotides is considered beneficial for epigenomic mapping of mammalian cells, and developed a novel MTase. More specifically, this novel MTase specifically recognizes CC dinucleotide sequences and can 5-methylate the cytosine residue at the 5' end of those sequences. This novel MTase, capable of recognizing and methylating CC dinucleotides—a mere two-base sequence that does not overlap with CG, where endogenous methylation occurs—has not been previously reported or realized.
[0020] 2. Protein The protein according to the present invention is a protein that functions as a DNA methyltransferase (MTase). Specifically, it is a protein that specifically recognizes CC dinucleotide sequences and has the activity to 5-methylate the cytosine residue at the 5' end of the sequence. More specifically, it is a protein that has acquired the above activity (the activity to 5-methylate the cytosine residue at the 5' end of the CC dinucleotide sequence) by deleting some amino acid residues, including the N-terminus, of the MTase of M.CviQIX or M.CviPII.
[0021] Here, M.CviQIX is an MTase that has been identified as originating from Paramecium bursaria Chlorella virus NY2A, and its amino acid sequence is the one shown in Sequence ID No. 2. While not limited to this, the base sequence encoding this amino acid sequence could be, for example, the base sequence shown in Sequence ID No. 1. M.CviPII is an MTase derived from Chlorella virus, and its amino acid sequence is the one shown in SEQ ID NO: 4. While not limited to this, the nucleotide sequence encoding this amino acid sequence could be, for example, the one shown in SEQ ID NO: 3.
[0022] Furthermore, the MTases of M.CviQIX and M.CviPII are both known to be among many conventionally known MTases, specifically capable of recognizing the "CCD trinucleotide sequence" (where D is A, G, or C) and 5-methylating the cytosine residue at its 5' end. There is generally no rational reason for a person skilled in the art to further modify an MTase capable of recognizing and methylating such a relatively short sequence as the CCD trinucleotide. Even if one were to attempt to modify it into an MTase capable of recognizing and methylating an even shorter dinucleotide sequence, there would be no rational reason to resort to deleting the amino acid residues including the N-terminus.
[0023] The protein in question is not limited to M.CviQIX, which is an MTase from which some amino acid residues, including the N-terminus, have been deleted. For example, it may be an MTase from which 11 amino acid residues, 13 amino acid residues, or 15 amino acid residues have been deleted from the N-terminus, but the one from which 15 amino acid residues have been deleted is preferred. The amino acid sequence of this protein from which 15 amino acid residues have been deleted from the N-terminus of M.CviQIX is the amino acid sequence shown in SEQ ID NO: 6 (the nucleotide sequence encoding this amino acid sequence is the nucleotide sequence shown in SEQ ID NO: 5). Therefore, a protein having the amino acid sequence shown in SEQ ID NO: 6 is also preferred as a protein according to the present invention.
[0024] The protein in question is not limited to M.CviPII, which is MTase, with some amino acid residues, including the N-terminus, deleted. For example, it may be a protein in which 11 amino acid residues, 13 amino acid residues, or 15 amino acid residues have been deleted from the N-terminus of the MTase, but a protein in which 15 amino acid residues have been deleted is preferred. The amino acid sequence of this protein in which 15 amino acid residues have been deleted from the N-terminus of M.CviPII is the amino acid sequence shown in SEQ ID NO: 8 (the base sequence encoding this amino acid sequence is the base sequence shown in SEQ ID NO: 7). Therefore, a protein having the amino acid sequence shown in SEQ ID NO: 8 is also preferred as a protein according to the present invention.
[0025] The proteins according to the present invention are not limited to these, and for example, proteins (so-called mutant proteins) that include an amino acid sequence in which one or more amino acids are deleted, substituted, or added in the amino acid sequence of the various proteins described above in this section, and that have the activity to specifically recognize a CC dinucleotide sequence and 5-methylate the cytosine residue on the 5' side of the sequence, are also preferably mentioned. Here, the above-mentioned "amino acid sequence in which one or several amino acids are deleted, substituted, or added" is preferably an amino acid sequence in which, for example, 1 to 10 amino acids, preferably 1 to several, 1 to 5, 1 to 4, 1 to 3, 1 to 2, or 1 amino acid is deleted, substituted, or added. The introduction of such mutations, such as deletions, substitutions, or additions, is carried out using a mutation introduction kit utilizing site-directed mutagenesis, such as GeneTailor. TMThis can be performed using a Site-Directed Mutagenesis System (Invitrogen) and a TaKaRa Site-Directed Mutagenesis System (Prime STAR® Mutagenesis Basal kit, Mutan®-Super Express Km, etc.: manufactured by Takara Bio). Furthermore, whether or not the above-mentioned deletion, substitution, or addition mutations have been introduced can be confirmed using various amino acid sequencing methods, as well as structural analysis methods such as X-ray and NMR.
[0026] Furthermore, the proteins according to the present invention are not limited to these, and for example, proteins (so-called mutant proteins) that contain an amino acid sequence having 80% or more identity (homology) with the amino acid sequences of the various proteins described above in this section, and that have the activity to specifically recognize the CC dinucleotide sequence and 5-methylate the cytosine residue on the 5' side of the sequence are also preferred. Here, it is more preferable that the above-mentioned identity is 85% or more, 90% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more. The aforementioned mutant proteins can also be genetically engineered using the gene that codes for the amino acid sequence of the protein.
[0027] In this invention, "specifically recognizing CC dinucleotide sequences" means that only dinucleotides consisting of two cytosine (C) bases can be recognized. "5-methylating the cytosine residue at the 5' end of the CC dinucleotide sequence" means that the carbon at position 5 of the cytosine residue at the 5' end of the recognized CC is methylated, resulting in 5-methylcytosine or 5-methylcytidine (5mC). Such methylation activity can be detected, measured, and evaluated by methods such as examining the presence or absence of DNA cleavage by HaeIII (which recognizes GGCC) or HpaII (which recognizes CCGG), whose cleavage activity is known to be inhibited by the presence of 5mC, or by bisulfite sequencing.
[0028] The protein according to the present invention may be formed by combining and linking peptides derived from natural products, or it may be obtained by artificial chemical synthesis, and is not limited to these. Proteins derived from natural products may be obtained directly from natural products by known recovery and purification methods, or they may be obtained by incorporating the gene encoding the protein into various expression vectors, etc., using known genetic engineering techniques, introducing them into cells, expressing them, and then obtaining the protein by known recovery and purification methods. Alternatively, commercially available kits, such as the reagent kit PROTEIOS, may be used. TM (Toyobo), TNT TM System (Promega), PG-Mate synthesis device TM The protein may be produced using a cell-free protein synthesis system with Toyobo (Toyobo) and RTS (Roche Diagnostics), etc., and obtained by known recovery and purification methods, but is not limited to these methods.
[0029] Furthermore, chemically synthesized proteins can be obtained using known protein synthesis methods. Examples of synthesis methods include the azide method, acid chloride method, acid anhydride method, mixed acid anhydride method, DCC method, activated ester method, carvoimidazole method, and redox method. Both solid-phase and liquid-phase synthesis methods can be applied. Commercially available protein synthesizers may also be used. After the synthesis reaction, the protein can be purified by combining it with known purification methods such as chromatography. In the present invention, derivatives of the various proteins described above, or in place thereof, may be included. The term "derivative" includes all derivatives that can be prepared from the protein, and examples include those in which some of the constituent amino acids are substituted with unnatural amino acids, or those in which some of the constituent amino acids (mainly their side chains) are chemically modified.
[0030] Furthermore, the present invention may include, together with or in lieu of, the various proteins and / or derivatives of the present invention described above, salts of the proteins and / or derivatives. Physiologically acceptable acid addition salts or basic salts are preferred as the salts. Examples of acid addition salts include salts with inorganic acids such as hydrochloric acid, phosphoric acid, hydrobromic acid, and sulfuric acid, or salts with organic acids such as acetic acid, formic acid, propionic acid, fumaric acid, maleic acid, succinic acid, tartaric acid, citric acid, malic acid, oxalic acid, benzoic acid, methanesulfonic acid, and benzenesulfonic acid. Examples of basic salts include salts with inorganic bases such as sodium hydroxide, potassium hydroxide, ammonium hydroxide, and magnesium hydroxide, or salts with organic bases such as caffeine, piperidine, trimethylamine, and pyridine.
[0031] The salt can be prepared using a suitable acid such as hydrochloric acid, or a suitable base such as sodium hydroxide. For example, it can be prepared by processing in water or in a liquid containing an inert, water-miscible organic solvent such as methanol, ethanol, or dioxane, using a standard protocol.
[0032] 2. Recombinant genes The genes encoding the proteins of the present invention described above are not limited to, but include, for example, genes containing the DNA of (a), (b), or (c) below. The genes containing these DNAs may consist only of these DNAs, or they may contain these DNAs in part and also include known base sequences necessary for gene expression (transcription promoters, SD sequences, Kozak sequences, terminators, etc.) or recognition sequences for various restriction enzyme sites, and are not limited to, such genes.
[0033] (a) DNA containing the nucleotide sequence shown in Sequence ID No. 5 or 7. (b) DNA that hybridizes under stringent conditions with DNA having a base sequence complementary to the DNA in (a) above, and that encodes a protein that specifically recognizes a CC dinucleotide sequence and has the activity to 5-methylate the cytosine residue at the 5' end of the sequence. (c) DNA that has 80% or more identity with the DNA in (a) above, and that encodes a protein that specifically recognizes the CC dinucleotide sequence and has the activity to 5-methylate the cytosine residue at the 5' end of the sequence.
[0034] In the DNA described in (a) above, the nucleotide sequence shown in SEQ ID NO: 5 encodes the amino acid sequence obtained by deleting 15 residues from the N-terminus of the amino acid sequence of M.CviQIX MTase (SEQ ID NO: 2), as previously mentioned. Furthermore, the nucleotide sequence shown in SEQ ID NO: 7 encodes the amino acid sequence obtained by deleting 15 residues from the N-terminus of the amino acid sequence of M.CviPII MTase (SEQ ID NO: 4).
[0035] The DNA described in (b) above can be obtained from a cDNA library or genome library by performing known hybridization methods such as colony hybridization, plaque hybridization, and Southern blotting, using the DNA described in (a) above, DNA consisting of a complementary base sequence thereto, or fragments thereof as probes. The library may be one prepared by known methods, or a commercially available cDNA library or genome library may be used, and is not limited to these. For detailed procedures regarding hybridization, refer to Molecular Cloning, A Laboratory Manual 4th ed., Cold Spring Harbor Laboratory Press (1989), etc.
[0036] "Stringent conditions" in the hybridization method refer to the conditions during washing after hybridization, for example, a buffer salt concentration of 15-330 mM, a temperature of 25-65°C, preferably a salt concentration of 15-150 mM and a temperature of 45-55°C. Specifically, conditions such as 80 mM and 50°C can be cited. Furthermore, in addition to conditions such as salt concentration and temperature, various conditions such as probe concentration, probe length, and reaction time can also be considered to appropriately set the conditions for obtaining the DNA described in (b) above.
[0037] The DNA in (c) above is DNA that has 80% or more identity (homology) with the DNA in (a) above, and more preferably DNA that has 85% or more identity, 90% or more identity, 95% or more identity, 96% or more identity, 97% or more identity, 98% or more identity, or 99% or more identity. This DNA is also preferred as the DNA to be hybridized in (b) above. Mutant DNA, such as the DNA described in (c) above, can be prepared according to site-directed mutation induction methods described in, for example, Molecular Cloning, A Laboratory Manual 4th ed., Cold Spring Harbor Laboratory Press (1989), Current Protocols in Molecular Biology, John Wiley & Sons (1987-1997), etc. Specifically, it can be prepared using a mutation introduction kit utilizing site-directed mutagenesis, using known methods such as the Kunkel method or the Gapped duplex method. An example of such a kit is QuickChange. TM Site-Directed Mutagenesis Kit (manufactured by Stratagene), GeneTailor TM Preferred examples include the Site-Directed Mutagenesis System (manufactured by Invitrogen) and the TaKaRa Site-Directed Mutagenesis System (Mutan-K, Mutan-Super Express Km, etc.: manufactured by Takara Bio).
[0038] Alternatively, MTG can be prepared by using PCR primers designed to introduce missense mutations so that the bases represent the codons of the desired amino acids, and performing PCR under appropriate conditions using DNA containing the base sequence encoding wild-type MTG as a template. The DNA polymerase used for PCR is not limited, but it is preferably a highly accurate DNA polymerase, such as Pwo DNA (PolymerzeroS Diagnostics), Pfu DNA polymerase (Promega), Platinum Pfx DNA polymerase (Invitrogen), KOD DNA polymerase (Toyobo), KOD-plus-polymerase (Toyobo), etc. The reaction conditions for PCR can be set appropriately depending on the optimal temperature of the DNA polymerase used, the length and type of DNA to be synthesized, etc. However, for example, if the cycle conditions are as follows, it is preferable to perform a total of 20 to 200 cycles, with one cycle consisting of "90-98°C for 5-30 seconds (thermal denaturation and dissociation) → 50-65°C for 5-30 seconds (annealing) → 65-80°C for 30-1200 seconds (synthesis and extension)".
[0039] As for the DNA in (b) and (c) above, DNA consisting of a base sequence that is not completely identical in base sequence to the DNA in (a) above, but is completely identical in amino acid sequence after translation (i.e., DNA in which silent mutations have been induced in the DNA in (a) above) is particularly preferred. Regarding the description of the "activity that specifically recognizes the CC dinucleotide sequence and 5-methylates the cytosine residue at the 5' end of the sequence" in the DNA described in (b) and (c) above, the contents described in the previous section may be applied as appropriate.
[0040] The gene encoding the protein according to the present invention is not particularly limited in that the codons corresponding to individual amino acids after translation are included. It may include DNA that shows codons commonly used in mammals such as humans after transcription (preferably codons with high frequency of use), or it may include DNA that shows codons commonly used in microorganisms such as E. coli and yeast, or in plants (preferably codons with high frequency of use). Furthermore, if the protein encoding the protein according to the present invention includes the linker sequence described above, the gene may also include DNA encoding the amino acid sequence of that linker sequence.
[0041] 3. Recombinant vectors and transformants To express the protein according to the present invention, it is first necessary to construct a recombinant vector by incorporating the gene according to the present invention described above into an expression vector. In this case, the gene to be incorporated into the expression vector may, if necessary, have a transcription promoter, SD sequence (if the host is a prokaryotic cell), and Kozak sequence (if the host is a eukaryotic cell) ligated upstream, or a terminator ligated downstream, or other elements such as enhancers, splicing signals, poly-A addition signals, and selection markers ligated. Note that the elements necessary for gene expression, such as the transcription promoter, may be included in the gene from the beginning, or they may be used if they are originally included in the expression vector, and the manner in which each element is used is not particularly limited.
[0042] Various methods utilizing known genetic recombination techniques can be employed to incorporate the gene into the expression vector, such as methods using restriction enzymes or methods using topoisomerases. Furthermore, the expression vector is not limited to those that can hold the gene encoding the protein of the present invention, such as plasmid DNA, bacteriophage DNA, retrotransposon DNA, retroviral vectors, and artificial chromosome DNA, and a vector suitable for the host cell being used can be appropriately selected and used. Next, the constructed recombinant vector is introduced into a host to obtain a transformant, which can then be cultured to express the protein of the present invention. In this invention, "transformant" refers to a host into which a foreign gene has been introduced. This includes, for example, a host into which a foreign gene has been introduced by introducing plasmid DNA or the like (transformation), and a host into which a foreign gene has been introduced by infecting the host with various viruses and phages (transduction).
[0043] The host is not limited to any host capable of expressing the protein of the present invention after the recombinant vector has been introduced, and can be appropriately selected. Examples of known hosts include various animal cells such as human and mouse cells, various plant cells, bacteria, yeast, and other plant cells. When using animal cells as hosts, examples include human fibroblasts, CHO cells, monkey cells COS-7 and Vero, mouse L cells, rat GH3 cells, and human FL cells. Insect cells such as Sf9 cells and Sf21 cells can also be used. When using bacteria as a host, for example, E. coli and Bacillus subtilis are used. When using yeast as a host, for example, Saccharomyces cerevisiae and Schizosaccharomyces pombe are used. When using plant cells as a host, for example, tobacco BY-2 cells are used.
[0044] The method for obtaining transformants is not limited and can be appropriately selected considering the combination of host and expression vector types. However, preferred methods include electroporation, lipofection, heat shock, PEG, calcium phosphate, DEAE dextran, and infection with various viruses such as DNA viruses and RNA viruses. In the resulting transformants, the codon types of the genes contained in the recombinant vector may or may not match the codon types of the host actually used; they are not limited to these.
[0045] 4. Protein production methods The protein according to the present invention can be produced by a method comprising the steps of culturing the transformant described above and collecting from the resulting culture a protein having the activity to specifically recognize a CC dinucleotide sequence and 5-methylate the cytosine residue on the 5' side of the sequence. Here, "culture" means any of the following: culture supernatant, cultured cells, cultured bacterial cells, or cell or bacterial cell lysates. The culturing of the transformant can be carried out according to the usual method used for culturing the host. The target protein is accumulated in the culture. In the present invention, the collection step may include a protein purification step.
[0046] As for the culture medium used in the above cultivation, any known natural culture medium or synthetic culture medium may be used, as long as it contains a carbon source, nitrogen source, inorganic salts, etc. that the host can utilize, and is capable of efficiently culturing the transformants. During culture, selective pressure may be applied to prevent the loss of recombinant vectors and genes encoding the target protein in the transformants. That is, if the selection marker is a drug resistance gene, the corresponding drug can be added to the culture medium, and if the selection marker is a nutrient complement gene, the corresponding nutrient can be removed from the culture medium.
[0047] When culturing transformants transformed with an expression vector using an inducible promoter, a suitable inducer (e.g., IPTG) may be added to the culture medium as needed. The culture conditions for the transformants are not particularly limited as long as they do not hinder the productivity of the target protein and the growth of the host, and are usually carried out at 10°C to 40°C, preferably 20°C to 37°C, for 5 to 100 hours. pH can be adjusted using inorganic or organic acids, alkaline solutions, etc. Culture methods include solid culture, static culture, shaking culture, and aerated stirring culture.
[0048] If the target protein is produced within the bacterial cells or cells after culturing, the target protein can be collected by disrupting the bacterial cells or cells. Methods for disrupting bacterial cells or cells include high-pressure treatment using a French press or homogenizer, sonication, grinding with glass beads, enzymatic treatment using lysozyme, cellulase, or pectinase, freeze-thaw treatment, hypotonic solution treatment, and lysis induction treatment using phages. After disruption, the disruption residue (including the cell extract insoluble fraction) can be removed as needed. Methods for removing the residue include centrifugation and filtration, and if necessary, flocculants or filter aids can be used to increase the efficiency of residue removal. The supernatant obtained after removing the residue is the cell extract soluble fraction and can be used as a crudely purified protein solution.
[0049] Furthermore, if the target protein is produced within or inside the bacterial cells, the bacterial cells or cells themselves can be recovered by centrifugation, membrane separation, etc., and used without being disrupted. On the other hand, if the target protein is produced outside the bacterial cells or cells, the culture medium can be used as is, or the bacterial cells or cells can be removed by centrifugation or filtration. Subsequently, the target protein can be collected from the culture by extraction using ammonium sulfate precipitation or other methods as needed, and further isolation and purification can be performed using dialysis or various chromatography methods (gel filtration, ion exchange chromatography, affinity chromatography, etc.) as necessary.
[0050] The production yield of proteins obtained by culturing transformants can be confirmed, for example, per culture medium, per wet or dry cell weight, or per crude enzyme solution protein, by SDS-PAGE (polyacrylamide gel electrophoresis). Furthermore, the production of the target protein can be carried out not only using the protein synthesis system with transformants described above, but also using a cell-free protein synthesis system that does not use any living cells.
[0051] A cell-free protein synthesis system is a system that synthesizes a target protein in an artificial container such as a test tube using a cell extract. In addition, the cell-free protein synthesis systems that can be used include cell-free transcription systems that synthesize RNA using DNA as a template. In this case, the origin of the cell extract used is preferably the host cell described above. As the cell extract, for example, extracts derived from eukaryotic cells or prokaryotic cells can be used. More specifically, extracts of CHO cells, rabbit reticulocytes, mouse L-cells, HeLa cells, wheat germ, budding yeast, Escherichia coli, etc. can be used. These cell extracts may be used after concentration or dilution, or may be used as they are, and are not limited.
[0052] The cell extract can be obtained, for example, by ultrafiltration, dialysis, polyethylene glycol (PEG) precipitation, etc. Such cell-free protein synthesis can also be carried out using commercially available kits. For example, the reagent kit PROTEIOS TM (Toyobo), TNT TM System (Promega), the synthesis device PG-Mate TM (Toyobo), RTS (Roche Diagnostics), etc. can be mentioned. The target protein produced by cell-free protein synthesis can be purified by appropriately selecting means such as chromatography as described above.
[0053] 5. Use of the protein An epigenome analysis method (epigenome mapping) including using the protein according to the present invention (including the protein obtained by the production method in the previous section) is also included in the present invention.
[0054] Despite the increasing demand for MTases as tools for epigenome mapping, the available enzymes are limited, particularly for cytosine C-5 MTases with short recognition sequences. The protein according to the present invention is capable of recognizing and methylating CC dinucleotide sequences, which are different from (and do not overlap with) the CG dinucleotide sequences that undergo endogenous methylation in mammalian cells. Therefore, the protein according to the present invention is useful as a novel probe for epigenome analysis methods (epigenome mapping).
[0055] The present invention will be described more specifically below with reference to examples, but the present invention is not limited to these examples. [Examples]
[0056] 1. Materials and Methods Sequence analysis The amino acid sequence for M. CviQIX was obtained from REBASE (http: / / rebase.neb.com / rebase / rebase.html) (references 19, 20), and the amino acid sequence for M. CviPII was obtained from reference 21.
[0057] Artificial gene synthesis and subcloning into expression vectors The genes encoding M.CviQIX and M.CviPII were synthesized by Eurofin Genomics Inc. after codon optimization of E. coli. Genes were synthesized that did not contain the recognition sequences for BamHI and EcoRI internally, but had the recognition sequences for BamHI and EcoRI at their 5' and 3' ends, respectively (SEQ ID NOs. 9 and 10 (the nucleotide sequences of SEQ ID NOs. 9 and 10 are obtained by adding the above recognition sequences to the 5' and 3' ends of the nucleotide sequences of SEQ ID NOs. 5 and 7, respectively)). These gene fragments were subcloned into the BamHI-EcoRI region of pCZS (Figure 1). The creation of N-terminal deletion mutant genes for M.CviQIX and M.CviPII was achieved by performing inverse PCR on expression vectors containing their full-length versions, followed by circularization via infusion cloning (Takara Bio Inc.).
[0058] Small-scale induction of MTase and bacterial genomic DNA extraction Bacterial expression vectors pCZS containing genes encoding M.CviQIX and M.CviPII were introduced into T7Express (New England Biolab, Ipswich, MA). These cells were inoculated with 100 μg / mL of carbenicillin (Nacalai Tesque, Kyoto, Japan) in 3 mL of LB medium (1% [w / v] tryptone, 0.5% [w / v] yeast extract, and 0.5% [w / v] sodium chloride) and grown overnight at 37°C with vigorous shaking. 50 microliters of each cell suspension were diluted in 3 mL of carbenicillin-supplemented LB medium, and the cells were cultured at 37°C for 3 hours with shaking. The cell suspension was cooled in ice water for 30 minutes, and protein expression was induced by adding 3 μL of 1 M isopropyl β-D-1 thiogalactopyranoside (IPTG, Nacalai Tesque). The cells were collected by centrifugation and used for genomic DNA purification using the DNeasy Blood & Tissue Kit (Qiagen). Genomic DNA concentrations were measured using the Qubit dsDNA BR assay kit and Qubit fluorometer (Thermo Fisher Scientific, Waltham, MA, USA).
[0059] Analysis of DNA methylation by restriction enzyme digestion Genomic DNA (100 ng) was added to 10 μL of a reaction mixture containing 1×CutSmart Buffer and restriction enzyme HaeIII or HpaII. The reaction was incubated at 37°C for 1 hour, and the enzymes were thermally inactivated at 70°C for 10 minutes. After treatment with one of the restriction enzymes, the DNA was analyzed using the E-Gel Power SNAP system (Thermo Fisher Scientific) on a 2% E-Gel Ex gel.
[0060] Purification of recombinant MTase Bacterial cells transformed with pCZS expressing the gene encoding MTase were inoculated into 30 mL of LB medium containing 100 μg / mL carbenicillin and cultured overnight at 37°C with vigorous shaking. Seed cultures were transferred to a 2 L Erlenmeyer flask with baffles containing 1 L of LB medium supplemented with carbenicillin and cultured further at 37°C with shaking for 4 hours. This flask was cooled in ice water for 30 minutes, and then protein expression was induced by adding 230 mg of IPTG. The flask was incubated overnight at 16°C with vigorous shaking. Cells were collected by centrifugation at 2500 × g for 15 minutes, and the collected cells were frozen at -80°C until use.
[0061] Bacterial cells were resuspended in 20 mL of HisTrap buffer A (20 mM sodium phosphate, pH 7.5, 800 mM NaCl, 20 mM imidazole, 1 mM DTT, and 10% glycerol) containing a protease inhibitor cocktail (Nacalai Tesque), and disrupted by sonication. The lysate was removed by centrifugation at 15,000 × g for 15 minutes and filtered through a 0.45 μm syringe filter. The first affinity purification was performed using HisTrap buffer A and HisTrap buffer B (20 mM sodium phosphate, pH 7.5, 800 mM NaCl, 200 mM imidazole, 1 mM DTT, 10% glycerol) with a 5 mL HisTrap HP column (Cytiva, Marlborough, MA). The fraction containing the target protein was further purified using a 1 mL StrepTrap HP (Cytiva) with StrepTrap Buffer A (100 mM Tris-HCl, pH 8.0, 150 mM NaCl, 1 mM EDTA, 1 mM DTT) and StrepTrap Buffer B (0 mM Tris-HCl, pH 8.0, 150 mM NaCl, 1 mM EDTA, 1 mM DTT, and 2.5 mM d-desthiobiotin). Chromatographic purification was performed using an AKTA Start System (Cytiva). The dual-affinity purified protein was concentrated by ultrafiltration using an Amicon Ultra-4 30 K instrument (Millipore), and bovine serum albumin (BSA) and glycerol were added to obtain final concentrations of 100 μg / mL and 50% (v / v), respectively. The solution was then stored at -20°C until use.
[0062] Budding yeast NOMe-Seq S288C cells (Biological Resources Center, National Institute of Technology and Evaluation, NBRC1136) were cultured in 50 ml of YPD medium (1% yeast extract, 2% bactopeptone, 2% dextrose) and incubated overnight at 30°C with shaking. The cell suspension was diluted with 1 L of YPD and incubated again at 30°C with shaking for 4 hours. The cells were collected by centrifugation at 2000 × g for 15 minutes, washed twice with 50 mL of water, and resuspended in 20 mL of 1 M sorbitol. 20 μL of β-marcaptoethanol and 20 mg of Zymolyase 100T (Seikagaku Kogyo Co., Ltd., Japan) were added to the cell suspension and incubated at room temperature for 15 minutes to digest and remove the cell wall and obtain spheroplasts. Next, the spheroplasts were washed twice with 20 mL of 1 M sorbitol and resuspended in 20 mL of 1 M sorbitol. 240 μL of 38% formaldehyde (Wako Chemical, Osaka, Japan) was added to the resuspended spheroplasts, and they were fixed by incubation at room temperature for 5 minutes. The fixed spheroplasts were washed twice with 20 mL of 1 M sorbitol and pelletized by centrifugation at 5000 × g for 5 minutes. The spheroplasts were then resuspended in 10 mL of nucleus preparation solution (10 mM HEPES-KOH, pH 7.5, 10 mM NaCl, 3 mM MgCl2, and 0.5% [v / v] Nonidet P 40 substitute) and incubated at room temperature for 5 minutes. The nuclei were collected by centrifugation at 5000 × g for 5 minutes and washed twice with 10 mL of nucleus washing solution (10 mM HEPES-KOH, pH 7.5, 10 mM NaCl, 3 mM MgCl2). The nuclei were resuspended in 10 mL of nuclear washing solution, and 500 μL of this solution was dispensed into 20 tubes, which were stored at -80°C until use.
[0063] The frozen nuclear suspension was thawed and resuspended in 500 μL of 1× methylated buffer (50 mM Tris-HCl, pH 8.5, 50 mM NaCl, 1 mM DTT, 1.6 mM SAM). Next, an appropriate amount of purified M.CviQIX N-terminal 15 amino acid residue-deficient protein (referred to as CCMT; see the results section below) was added. After incubation at 37°C for 1 hour, the nuclei were pelletized by centrifugation at 5000×g for 1 minute and resuspended in 40 μL of 10 mM Tris-HCl, pH 8.0. 20 mg / ml proteinase K and 5 μL of 10% (w / v) SDS were added to this nuclear suspension and incubated at 50°C for 1 hour. Next, 150 μL of 1 M TrisHCl, pH 8.0 was added to this reaction, the mixture was topped with white mineral oil, and incubated at 80°C for 24 hours. The aqueous phase was transferred to a new test tube and mixed with 200 μL of Buffer AL and isopropanol. This mixture was loaded onto a DNeasy column by centrifugation at 10,000 × g for 1 minute, and then sequentially washed with 500 μL each of Buffer W1 and Buffer W2. Finally, the purified DNA was eluted with 200 μL of Buffer AE and used for WGBS analysis using the tPBAT protocol. Buffer AL, Buffer W1, Buffer W2, and Buffer AE are reagents included in the DNeasy Blood & Tissue Kit.
[0064] NOMe-Seq in mammalian cells IMR-90 cells (JCRB Cell Bank #JCRB9054) were cultured at 37°C in 100% humidity and 5% CO2 atmosphere in four T75 flasks containing 30 mL of Minimum Essential Medium (Thermo Fisher Scientific) supplemented with 10% fetal bovine serum (Thermo Fisher Scientific) and 50 U / mL penicillin-streptomycin (Thermo Fisher Scientific). The grown cells were treated with trypsin and collected by centrifugation at 2000 × g for 5 minutes. The cells were washed twice with 20 mL of D-PBS (Nacalai Tesque) and resuspended in CELLBANKER 1 (Takara Bio) at 1 × 10⁷ cells / ml. The cell suspension was aliquoted in 500 μL and stored at -80°C until use.
[0065] Before use, the frozen cell suspension was thawed and washed twice with PBS. Next, the cells were resuspended in nuclear preparation buffer (10 mM Tris-HCl, pH 8.0, 10 mM NaCl, 3 mM MgCl2, 0.1 mM EDTA) and 0.5% (w / v) NP-40 and incubated on ice for 10 minutes. Then, the nuclei were collected by centrifugation at 2000 × g for 5 minutes, resuspended in 500 μL of 1 × methylated buffer in an optimal volume of CCMT, and incubated at 37°C for 30 minutes. Subsequently, the nuclei were collected by centrifugation at 2000 × g for 5 minutes and the supernatant was removed. The nuclei were dissolved in 20 μL of 20 mg / ml proteinase K and 200 μL of Buffer ATL was added. The reaction mixture was incubated at 56°C for 1 hour, after which 200 μL of Buffer AL and 200 μL of isopropanol were added. This mixture was loaded onto a DNeasy column and washed with 500 μL of Buffer W1 or Buffer W2. The purified DNA was eluted with 200 μL of Buffer AE and used for WGBS analysis.
[0066] 2.Results N-terminal deletion enhances MTase activity of M.CviQIX and M.CviPII. The inventors attempted to produce an active protein of M.CviQIX, which was described in databases as having MTase activity against the first C of CCD trinucleotides. However, DNA methylation activity could not be detected in most bacterial cells with several expression vectors from which this gene was subcloned (Figure 3A-3C). However, relatively strong MTase activity was observed in cells expressing M.CviQIX subcloned with expression vectors pCTF and pCNS. In other words, the genomic DNA of cells with these two types of expression vectors showed resistance to both HaeIII and HpaII digestion (Figure 3C and Figures 7A-7B). pCTF and pCNS are designed to produce proteins fused with a trigger factor (TF) and SUMO at the N-terminus of M.CviQIX, respectively (Figure 3A). TF and SUMO are expected to have effects on protein solubilization and folding (Reference 22). Furthermore, genomic DNA extracted from bacterial cells containing an expression vector with a SUMO-labeled N-terminus showed greater resistance to treatment with methylation-sensitive restriction enzymes compared to genomic DNA extracted from bacterial cells containing an expression vector with a SUMO-labeled C-terminus (Figure 3C). Therefore, we attempted to purify the products of these N-terminal TFs and SUMO-labeled constructs. However, after affinity purification using an N-terminal His6 tag, both proteins showed almost no MTase activity. These results suggested that some factor other than improved protein solubility or folding may be at the N-terminus and may be controlling the DNA methylation activity of M.CviQIX. Therefore, we hypothesized that the N-terminal sequence might be the cause of the low methylation activity of purified M.CviQIX. Accordingly, we decided to analyze the genomic DNA of bacterial cells containing an expression vector with an N-terminal knockout protein of M.CviQIX. The N-terminus of M.CviQIX has three methionine residues. Therefore, we hypothesized that in the cells of organisms that originally possess this gene, one of these three residues might be used for translation initiation. Therefore, we created three types of mutants in which the amino acid residue immediately preceding each methionine residue was deleted, and analyzed the methylation status of their host genomes.As a result, deletions of the N-terminal amino acids 11, 13, or 15 of M.CviQIX were found to enhance host genome methylation (Figures 3D-3E and 7C). In other words, any of these three methionine residues could be used as start codons to produce active M.CviQIX. Similar results were obtained for M.CviPII (Figures 7F-7G), indicating that the N-terminal amino acid sequences of M.CviQIX and M.CviPII exhibit inhibitory effects on DNA methylation activity.
[0067] Determination of recognition sequences for M.CviQIX and M.CviPII using WGBS. The inventors induced the expression of N-terminal 15-amino acid deletion (NDEL15) mutants of M.CviQIX and M.CviPII in E. coli cells, extracted bacterial genomic DNA, and performed whole-genome bisulfite sequencing (WGBS) (Figure 4A). Methylation levels were measured for each cytosine residue, and the frequency of occurrence of surrounding nucleotide sequences contributing to the methylation level was displayed as a logo. The results revealed that M.CviQIX and M.CviPII introduce methylation to the first C of the CC dinucleotide.
[0068] Purification of MTase and confirmation of its activity in vitro Next, the inventors attempted to purify active proteins using M.CviQIX and M.CviPII protein genes lacking the N-terminal 15 amino acid residues. Simultaneously, purification of the active protein of M.CviPI, for which purification conditions have been established, was attempted as a control. Bacterial cells containing expression vectors containing these genes were cultured, protein expression was induced, and these MTases were purified using a dual affinity purification method with HisTag and StrepII-tag (Figure 2). The proteins were stored in the presence of 100 μg / ml BSA, 10 mM DTT, and 50% glycerol, and enzyme activity was preserved during storage at -20°C. The purified proteins appeared highly homogeneous on SDS-PAGE. Assays with restriction enzymes and WGBS using λ phage DNA as a substrate confirmed that these MTases exhibited methylation activity, as they inhibited DNA cleavage (Figure 8B).
[0069] NOMe-Seq, which possesses CC-recognizing MTase, generates a methylation pattern similar to that of GCMT. CC dinucleotides have short recognition sequences, which can enable fine epigenome mapping with a resolution of approximately 10 bp. Therefore, the inventors attempted to establish M.CviQIX or M.CviPII, MTases that recognize CC dinucleotides, as probes for epigenome mapping. After comparing expression vector and gene combinations, M.CviQIX expressed under the control of a cold shock promoter showed the highest yield and activity, and this protein was named CCMT. By investigating the salt concentration in the methylation reaction of CCMT, it was found that 50 mM sodium chloride was the optimal condition for precise DNA methylation activity (Figure 9).
[0070] Next, CCMT was used for NOMe-Seq analysis of formaldehyde-bridged budding yeast nuclei. Formalin-fixed yeast spheroplasts were permeabilized in a buffer containing NP-40 to remove the cell membrane, and then methylation of DNA on the chromatin was performed by applying a methylation reaction with CCMT in the presence of S-adenosylmethionine. After purifying the genomic DNA, the methylation status of this DNA was analyzed by WGBS.
[0071] By calculating the methylation levels of 10-bp bins and comparing the DNA methylation status of nuclei treated with the commercially available M.CviPI (GCMT), which recognizes GC dinucleotides, and nuclei treated with CCMT, the two datasets showed similar genomic distribution patterns for 5mC. Since budding yeast lacks endogenous DNA methylation, all detected D5mC can be interpreted as having been introduced by either GCMT or CCMT. Similar trends were observed in the genomic DNA methylation patterns, as shown in Figures 5A-5B. Agglutination plots of CCMT and GCMT signals centered on several similar genomic features (Figures 5C-5D and 10A). Sharp DNA methylation peaks were observed in the promoter regions of genes (Figure 5B). These strong DNA methylation signals co-localized well with signals detected using ATAC-Seq. Therefore, these signals were considered to indicate open chromatin regions. Peaks detected by CCMT and GCMT-based NOMe-Seq and ATAC-Seq co-localized well (Figure 10B).
[0072] DNA methylation levels in CC and GC dinucleotides were well aggregated in the promoter-proximal region (Figure 5C-5D). These signals are undoubtedly from nucleosome-free regions, frequently observed in gene promoters as areas with little or no nucleosome binding. Promoter-proximal DNA methylation signals in the -100 to -250 region of the transcription start site were inversely correlated with low-signal regions detected via MNase-Seq (Figure 5E). The DNA methylation signals observed at the transcription start site showed a periodicity of 160-170 bp and were inversely correlated with signals observed by MNase-Seq (Figure 5B, 5C-5E). These observations reproduced the characteristics of previously reported NOMe-Seq and MNase-Seq (Reference 23).
[0073] CCMT-based NOMe-Seq achieves fine mapping of the mammalian epigenome. Next, CCMT-based NOMe-Seq was applied to IMR-90 cells, a human-derived cultured cell line. After permeabilization of the cell membrane with a buffer containing NP-40, CCMT treatment was performed, and genomic DNA was purified and subjected to WGBS analysis. Samples without CCMT treatment were also prepared and used as a control. High-quality methylome data with coverage 25.7-fold and 24.5-fold, respectively, were obtained for the samples with and without CCMT treatment (Table 1).
[0074] [Table 1]
[0075] As expected, the intrinsic DNA methylation levels at the CG site were almost identical between CCMT-treated and untreated cells (Figure 6A-6B). On the other hand, DNA methylation levels at the CC site increased only in CCMT-treated nuclei (Figure 6C-6D). These results demonstrate that endogenous CG methylation and CCMT-induced methylation at the CC dinucleotide can be explicitly separated.
[0076] Next, the two datasets were visually compared using a genome browser. As shown in Figure 6E, several large domains with distinct DNA methylation levels were observed in both datasets. These domains were undoubtedly highly methylated and moderately methylated regions, their presence described for IMR-90 cells (reference 24). CCMT-based NOMe comparative Seq and untreated control showed that the endogenous methylation state of mammalian cells targeting CG dinucleotides was observed to be nearly identical in both datasets. On the other hand, strong, sharp methylation peaks of CCMT methylation were observed in CCMT-treated samples, while they were not observed in the untreated control (Figure 6E). Furthermore, these peaks largely co-localized with peaks identified using ATAC-Seq. Therefore, the peaks observed in the CCMT-induced methylation signal indicated highly accessible chromatin regions (Figure 6E). These results strongly support the idea that CCMT-based NOMe-Seq can simultaneously measure both the intrinsic methylome and chromatin accessibility of mammalian cells.
[0077] 3. Discussion Despite the increasing demand for MTases as tools for epigenome mapping, the availability of enzymes is limited, particularly for cytosine C-5 MTases targeting short recognition sequences. In this study, we revealed that the methylation activity of M.CviQIX and M.CviPII is inhibited by the amino acid sequence of the first few residues at their N-terminus, and conversely, their N-terminal knockout strains exhibit high DNA methylation activity. These artificial mutant proteins were purifiable, and their activity was maintained in vitro. Furthermore, we demonstrated that the CC dinucleotide, a novel recognition sequence, can be targeted as a cytosine C-5 MTase. Since this CC dinucleotide never overlaps with the CG dinucleotide, which is the recognition sequence for endogenous mammalian MTases, it was considered to have a significant advantage as a probe for epigenetic mapping to simultaneously detect arbitrary epigenomic states in addition to endogenous mammalian MTases. In this example, this was demonstrated by the implementation of NOMe-Seq.
[0078] The following is a list of the references included in this specification. 1. Greenberg, MVC and Bourc'his, D. (2019) The diverse roles of DNA methylation in mammalian development and disease. Nature Reviews Molecular Cell Biology, 20, 590‐607. 2. Edwards, JR, Yarychkivska, O., Boulard, M. and Bestor, TH (2017) DNA methylation and DNA methyltransferases. Epigenetics & Chromatin, 10, 23. 3. Jones, PA (2012) Functions of DNA methylation: islands, start sites, gene bodies and beyond. Nature Reviews Genetics, 13, 484‐492. 4. Petryk, N., Bultmann, S., Bartke, T. and Defossez, P.‐A. (2021) Staying true to yourself: mechanisms of DNA methylation maintenance in mammals. Nucleic Acids Res., 49, 3020‐3032. 5. Jones, P.A. and Liang, G. (2009) Rethinking how DNA methylation patterns are maintained. Nat Rev Genet, 10, 805‐811.
[0079] 6. Edwards, J.R., Yarychkivska, O., Boulard, M. and Bestor, T.H. (2017) DNA methylation and DNA methyltransferases. Epigenetics Chromatin, 10, 23. 7. Tajima, S., Suetake, I., Takeshita, K., Nakagawa, A. and Kimura, H. (2016) Domain Structure of the Dnmt1, Dnmt3a, and Dnmt3b DNA Methyltransferases. Adv. Exp. Med. Biol., 945, 63‐86. 8. Jang, H.S., Shin, W.J., Lee, J.E. and Do, J.T. (2017) CpG and Non‐CpG Methylation in Epigenetic Gene Regulation and Brain Function. Genes, 8, 148. 9. Patil, V., Ward, R.L. and Hesson, L.B. (2014) The evidence for functional non‐CpG methylation in mammalian cells. Epigenetics, 9, 823‐828. 10. van de Lagemaat, L.N., Flenley, M., Lynch, M.D., Garrick, D., Tomlinson, S.R., Kranc, K.R. and Vernimmen, D. (2018) CpG binding protein (CFP1) occupies open chromatin regions of active genes, including enhancers and non‐CpG islands. Epigenetics & Chromatin, 11, 59.
[0080] 11. Teissandier, A. and Bourc'his, D. (2017) Gene body DNA methylation conspires with H3K36me3 to preclude aberrant transcription. The EMBO Journal, 36, 1471‐1473. 12. Dhayalan, A., Rajavelu, A., Rathert, P., Tamas, R., Jurkowska, R.Z., Ragozin, S. and Jeltsch, A. (2010) The Dnmt3a PWWP domain reads histone 3 lysine 36 trimethylation and guides DNA methylation. The Journal of biological chemistry, 285, 26114‐26120. 13. Solomon, M.J., Larsen, P.L. and Varshavsky, A. (1988) Mapping protein‐DNA interactions in vivo with formaldehyde: evidence that histone H4 is retained on a highly transcribed gene. Cell, 53, 937‐947. 14. Johnson, D.S., Mortazavi, A., Myers, R.M. and Wold, B. (2007) Genome‐wide mapping of in vivo protein‐DNA interactions. Science, 316, 1497‐1502. 15. Kelly, T.K., Liu, Y., Lay, F.D., Liang, G., Berman, B.P. and Jones, P.A. (2012) Genome‐wide mapping of nucleosome positioning and DNA methylation within individual DNA molecules. Genome Res., 22, 2497‐2506.
[0081] 16. Lee, I., Razaghi, R., Gilpatrick, T., Molnar, M., Gershman, A., Sadowski, N., Sedlazeck, F.J., Hansen, K.D., Simpson, J.T. and Timp, W. (2020) Simultaneous profiling of chromatin accessibility and methylation on human cell lines with nanopore sequencing. Nat Methods, 17, 1191‐1199. 17. Stergachis, A.B., Debo, B.M., Haugen, E., Churchman, L.S. and Stamatoyannopoulos, J.A. (2020) Single‐molecule regulatory architectures captured by chromatin fiber sequencing. Science, 368, 1449‐1454. 18. Steensel, B.v. and Henikoff, S. (2000) Identification of in vivo DNA targets of chromatin proteins using tethered Dam methyltransferase. Nat. Biotechnol., 18, 424‐428. 19. Roberts, R.J. and Macelis, D. (1997) REBASE‐restriction enzymes and methylases. Nucleic Acids Res., 25, 248‐262. 20. Roberts, R.J., Vincze, T., Posfai, J. and Macelis, D. (2015) REBASE‐‐a database for DNA restriction and modification: enzymes, genes and genomes. Nucleic Acids Res., 43, D298‐D299.
[0082] 21. Xu, M., Kladde, M.P., Van Etten, J.L. and Simpson, R.T. (1998) Cloning, characterization and expression of the gene coding for a cytosine‐5‐DNA methyltransferase recognizing GpC. Nucleic Acids Res., 26, 3961‐3966. 22. Malakhov, M.P., Mattern, M.R., Malakhova, O.A., Drinker, M., Weeks, S.D. and Butt, T.R. (2004) SUMO fusions and SUMO‐specific protease for efficient expression and purification of proteins. Journal of Structural and Functional Genomics, 5, 75‐86. 23. Klemm, S.L., Shipony, Z. and Greenleaf, W.J. (2019) Chromatin accessibility and the regulatory epigenome. Nature Reviews Genetics, 20, 207‐220. 24. Lister, R., Pelizzola, M., Dowen, R.H., Hawkins, R.D., Hon, G., Tonti‐Filippini, J., Nery, J.R., Lee, L., Ye, Z., Ngo, Q.M. et al. (2009) Human DNA methylomes at base resolution show widespread epigenomic differences. Nature, 462, 315‐322. 25. Braberg, H., Echeverria, I., Bohn, S., Cimermancic, P., Shiver, A., Alexander, R., Xu, J., Shales, M., Dronamraju, R., Jiang, S. et al. (2020) Genetic interaction mapping informs integrative structure determination of protein complexes. Science, 370.
[0083] 26. Miura, F., Fujino, T., Kogashi, K., Shibata, Y., Miura, M., Isobe, H. and Ito, T. (2018) Triazole linking for preparation of a next-generation sequencing library from single-stranded DNA. Nucleic Acids Res., 46, e95. [Sequence Listing Free Text]
[0084] Sequence ID 5: Recombinant DNA Sequence ID 6: Recombinant Protein Sequence ID 7: Recombinant DNA Sequence ID 8: Recombinant Protein Sequence ID 9: Synthetic DNA Sequence ID 10: Synthetic DNA
Claims
1. A protein of any one of the following (a), (b), or (c): (a) A protein having the amino acid sequence shown in SEQ ID NO: 6 or 8; (b) A protein comprising an amino acid sequence in which one or several amino acids are deleted, substituted, or added in the amino acid sequence of the protein of (a) above, and having an activity of specifically recognizing a CC dinucleotide sequence and methylating the cytosine residue on the 5'-side in the sequence; (c) A protein comprising an amino acid sequence having 80% or more identity with the amino acid sequence of the protein of (a) above, and having an activity of specifically recognizing a CC dinucleotide sequence and methylating the cytosine residue on the 5'-side in the sequence.
2. A gene encoding the protein according to Claim 1.
3. A gene containing any one of the following (a), (b), or (c) DNAs: (a) A DNA containing the nucleotide sequence shown in SEQ ID NO: 5 or 7; (b) A DNA that hybridizes under stringent conditions with a DNA consisting of a nucleotide sequence complementary to the DNA of (a) above, and encoding a protein having an activity of specifically recognizing a CC dinucleotide sequence and methylating the cytosine residue on the 5'-side in the sequence; (c) A DNA having 80% or more identity with the DNA of (a) above, and encoding a protein having an activity of specifically recognizing a CC dinucleotide sequence and methylating the cytosine residue on the 5'-side in the sequence.
4. A recombinant vector containing the gene according to Claim 2 or 3.
5. A transformant containing the recombinant vector according to Claim 4.
6. A step of culturing the transformant according to Claim 5, and A step of collecting from the obtained culture a protein having an activity of specifically recognizing a CC dinucleotide sequence and methylating the cytosine residue on the 5'-side in the sequence A method for producing the protein, comprising the above steps.
7. A step of performing a reaction using an in vitro protein synthesis system containing the recombinant vector according to Claim 4, and A step of collecting from the obtained reaction solution a protein having an activity of specifically recognizing a CC dinucleotide sequence and methylating the cytosine residue on the 5'-side in the sequence A method for producing the protein, comprising the above steps.
8. (1) A protein of any one of the following (a), (b), or (c): (a) A DNA methyltransferase that specifically recognizes the CCD trinucleotide sequence of M.CviQIX or M.CviPII and has the activity of 5-methylating the cytosine residue on the 5'-side in the said sequence; (b) A protein comprising an amino acid sequence in which one or several amino acids are deleted, substituted or added in the amino acid sequence of the protein of (a) above, and that specifically recognizes the CCD trinucleotide sequence and has the activity of 5-methylating the cytosine residue on the 5'-side in the said sequence; (c) A protein comprising an amino acid sequence having 80% or more identity with the amino acid sequence of the protein of (a) or (b) above, and that specifically recognizes the CCD trinucleotide sequence and has the activity of 5-methylating the cytosine residue on the 5'-side in the said sequence; The step of preparing a polynucleotide sequence encoding a protein in which a part of the amino acids including the N-terminus of is deleted and that specifically recognizes the CC dinucleotide sequence and has the activity of 5-methylating the cytosine residue on the 5'-side in the said sequence, (2) The step of culturing a host cell into which a recombinant vector containing the said polynucleotide sequence has been introduced, or the step of performing a reaction using an in vitro protein synthesis system containing a recombinant vector containing the said polynucleotide sequence, and (3) The step of collecting from the culture solution or reaction solution obtained in the step of (2) above, a protein that specifically recognizes the CC dinucleotide sequence and has the activity of 5-methylating the cytosine residue on the 5'-side in the said sequence A method for producing the said protein, which includes the above steps.
9. An epigenome analysis method, which includes using the protein according to Claim 1 or the protein produced by the method according to Claim 8.
10. An epigenome analysis method, which includes using the protein produced by the method according to Claim 6.
11. An epigenome analysis method, which includes using the protein produced by the method according to Claim 7.