A method for predicting epigenetic modification signals on DNA sequences
Through the CNN model based on WGBS data, the DNA sequence and 5mC characteristics are integrated, and DNA methylation signals are directly extracted from WGBS data, solving the problem of predicting the 5hmC modification level of the whole genome in the prior art, achieving high-precision prediction, and reducing the dependence on biological detection data.
Patent Information
- Application Number
- CN202410783593.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-18
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-06-18
AI Technical Summary
The prior art is difficult to achieve 5hmC modification level prediction with single base resolution of the whole genome, and it relies on a variety of biological detection technologies and histone modification signals, resulting in inconsistent model training targets and accumulation of errors, and the prediction accuracy is not high.
A convolutional neural network (CNN) model based on WGBS data is proposed. By integrating DNA sequences and 5mC features, DNA methylation signals are directly extracted from WGBS data, and a single-base resolution prediction of the 5hmC modification level is achieved.
The 5hmC modification level prediction with single base resolution of the whole genome is achieved, with higher accuracy than predictions on interval length, reducing dependence on biological detection data, and making the prediction model more operable.
Smart Images

Figure CN118430662B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biological information processing, and in particular to a method for predicting epigenetic modification signals on a DNA sequence. Background Art
[0002] 5-Hydroxymethylcytosine (5hmC) is a common oxidized form of DNA methylation. It is stably present in the genome of organisms and plays an important regulatory role in gene expression and cellular processes. In order to achieve unbiased quantification of 5hmC with single-base resolution, it is usually converted by chemical or bioenzymatic methods and then subjected to high-throughput sequencing, such as oxBS-Seq and CAPS sequencing. The experiment requires meticulous handling of biological samples and detection technology, so a prediction algorithm based on machine learning is proposed.
[0003] In 2020, the DeepH&M algorithm was proposed. It uses restriction endonucleases to cut DNA fragments carrying 5hmC modifications. After antibody enrichment sequencing, it combines DNA sequences and uses deep learning algorithms to predict 5hmC modification levels at single-base resolution. DeepH&M is mainly composed of three modules: the CpG module accepts genomic and methylation feature inputs, the DNA module uses convolutional neural networks to process raw DNA sequence data, and the joint module combines the outputs of the CpG and DNA modules to predict 5hmC and 5mC at the same time. The CpG module of this method relies on features including MeDIP-seq, MRE-seq, and hmC-Seal, and relies on three biological detection technologies at the same time. It uses three modules to solve a complex task. The obvious disadvantages are: 1. The training objectives of each step are inconsistent and deviate from the macro goal. It is difficult for the trained model to achieve the optimal result; 2. Each step has errors. The error of the previous step will affect the result of the training of the next step. The accumulation of errors ultimately leads to poor final results.
[0004] In 2024, a new multimodal deep learning framework called Deep5hmC was introduced, which integrates DNA sequence information and histone modification signals to predict tissue / cell type-specific genome-wide 5hmC modifications. However, Deep5hmC can only predict modification levels at the regional level of 1000bp, and cannot achieve single-base resolution prediction. Summary of the invention
[0005] The object of the present invention is to achieve accurate prediction of 5hmC modification levels at single-base resolution across the whole genome. This study proposes an innovative 5hmC modification prediction algorithm based on WGBS. This algorithm does not rely on additional histone modification or 5hmC enrichment data. By only using common and easily accessible WGBS data, combining DNA sequence information, and adopting a convolutional neural network model, it realizes accurate prediction of 5hmC modification levels.
[0006] The object of the present invention is achieved by the following technical solutions:
[0007] A method for predicting epigenetic modification signals on a DNA sequence, the steps of the method include:
[0008] S1. Data preparation: Calculate the DNA methylation levels and 5hmC modification levels upstream and downstream of each CpG site for 5mC and 5hmC levels from WGBS and CAPS data respectively, and classify according to the calculated results to construct a training set and a test set respectively. The latter serves as the target data of the model, and classify according to chromosome numbers to construct a training set and a test set respectively;
[0009] S2. Feature representation: Represent according to the data features in the training set and the test set to construct the feature input and output required for training. Combine the DNA sequence of 50bp upstream and downstream and the methylation level data features of 50 CpG sites upstream and downstream in the training set and the test set for representation to construct the feature input matrix required for training;
[0010] S3. Model training and tuning: Use a convolutional neural network model and a mean square error loss function for training and parameter tuning to achieve prediction of 5hmC modification levels.
[0011] Further, in step S1, extract the 5hmC modification level on CpG from CAPS data as the target data, use the sites with chromosome numbers 1 and 2 as the test set, and the sites with chromosome numbers 3 and 4 as the training set.
[0012] Further, in step S2, from the WGBS data, extract the DNA methylation levels of 50 CpG upstream and downstream of each CpG, extract the DNA sequence of 50bp upstream and downstream of each CpG site, and encode the DNA sequence using the One-hot encoding method to form a 4×100 data matrix. Combine the DNA methylation levels extracted from WGBS to form a 5×100 data matrix, and its formula is as follows:
[0013]
[0014] Further, in step S3, the feature data matrix As the input of the convolutional neural network (CNN) model, the 5hmC modification level is used as the output. The CNN model is trained to learn the relationship between features and 5hmC. The model parameters are adjusted according to the prediction accuracy. The loss function used by the model during training is MSE, which is calculated as follows:
[0015]
[0016] Where n is the number of samples, is the true 5hmC modification level of the i-th sample, The model predicts the 5hmC modification level for the i-th sample.
[0017] Furthermore, in step S3, the trained CNN model is used to predict the test set data to evaluate the model's predictive ability for 5hmC modification levels.
[0018] The beneficial effects of the present invention are:
[0019] (1) The present invention proposes a CNN model to predict the 5hmC signal, which is difficult to detect, with single-base resolution across the entire genome, with higher accuracy than the prediction of 5hmC levels based on interval length;
[0020] (2) By integrating DNA sequence and 5mC features, DNA methylation and sequence features can be integrated. Compared with models that only use DNA sequences or rely only on histones detected by biological experiments, the information input of the model is expanded, making the model more comprehensive and accurate;
[0021] (3) It relies on DNA methylation signals determined by WGBS and does not require support from other omics data, thus reducing dependence on biological detection data and making the prediction model more operational;
[0022] (4) By adopting the CNN model and taking advantage of its local perception, the present invention can efficiently extract DNA sequence features, especially for 5hmC modification, thereby improving prediction accuracy and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 Build a flowchart for the algorithm;
[0024] Figure 2 This is the architecture diagram of the CNN model;
[0025] Figure 3 Schematic diagram of forward propagation. DETAILED DESCRIPTION
[0026] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0027] See also Figure 1 This embodiment provides a method for predicting epigenetic modification signals on a DNA sequence, the method comprising the steps of:
[0028] Data pre-processing: Extract 5mC and 5hmC modification levels upstream and downstream of CpG sites;
[0029] Feature representation processing: including one-hot encoding of DNA sequences to obtain DNA sequence matrix, and feature processing of 5mC;
[0030] Use the Caffe framework to train the model and optimize the parameters. During the model training process, Caffe calculates the loss function value based on the forward propagation, adjusts the network weights through the back-propagation algorithm, and automatically adjusts the model parameters through continuous iterative training to minimize the loss function. It also performs hyperparameter tuning based on the performance indicators on the validation set to optimize the model's prediction performance.
[0031] The 5hmC modification level on CpG was extracted from the CAPS data as the target data, the sites with chromosome numbers 1 and 2 were used as the test set, and the sites with chromosome numbers 3 and 4 were used as the training set.
[0032] The data set numbered GSM4708554 was downloaded from NCBI, and the DNA methylation levels of 50 CpGs upstream and downstream of each CpG were extracted from the WGBS data.
[0033] Extract 50 bp of DNA sequence upstream and downstream of each CpG site, and encode the DNA sequence using One-hot encoding to form 4 The data matrix of 100 was merged to form 5 The data matrix of 100 is as follows:
[0034]
[0035] The feature data matrix As the input of the convolutional neural network (CNN) model, the 5hmC modification level was used as the output. The CNN model was trained to learn the relationship between features and 5hmC. The model parameters were adjusted according to the prediction accuracy. The model structure is as follows: Figure 2As shown in the figure, it consists of two convolutional layers (32 convolution kernels of size 5×5 and 64 convolution kernels of size 3×3), two maximum pooling layers (both use pooling kernels of size 2×1), and three fully connected layers (containing 100, 100, and 1 neurons respectively).
[0036] See also Figure 3 In the forward propagation process, the 100×5×1 feature matrix of each DNA is output through the convolution layer, pooling layer and fully connected layer. The matrix first enters the first convolution layer to extract the local features of the data and capture different feature patterns through multi-channel output; the first pooling layer reduces the dimension of the data, retains important features, and prevents overfitting; the second convolution layer further convolves on the extracted primary features to obtain more complex patterns and features, and increases the number of feature mappings to capture more detailed information; the second pooling layer further reduces the amount of data, retains important feature information, and improves the generalization ability of the model; the three fully connected layers gradually linearly combine and map the extracted features to strengthen the association between features, and finally output the regression value, which represents the prediction of the 5hmC modification level.
[0037] The loss function used by the model during training is MSE, which is calculated as follows:
[0038]
[0039] The trained deep learning CNN model was used to predict the test set data and evaluate the model's ability to predict 5hmC modification levels.
[0040] The method proposes a prediction based on a CNN model for 5hmC signals that are difficult to detect, with single-base resolution for the entire genome, and the accuracy is higher than that of 5hmC level prediction on an interval length; by integrating DNA sequence and 5mC features, DNA methylation and sequence features can be fused, and compared with models such as histones that only use DNA sequences or only rely on biological experimental detection, the information input of the model is expanded, making the model more comprehensive and accurate; relying on DNA methylation signals determined by WGBS, no support from other omics data is required, thereby reducing dependence on biological detection data and making the prediction model more operational; adopting the CNN model and utilizing its local perception advantage, the present invention can efficiently extract DNA sequence features, especially for 5hmC modification, thereby significantly improving prediction accuracy and efficiency.
[0041] The above is only a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be modified within the scope of the concept described herein through the above teachings or the technology or knowledge of the relevant field. The changes and modifications made by those skilled in the art shall not deviate from the spirit and scope of the present invention, and shall be within the scope of protection of the claims attached to the present invention.
Claims
1. A method for predicting epigenetic modification signals on a DNA sequence, characterized in that: The method steps include: S1. Data preparation: 5mC and 5hmC modification levels of each CpG site upstream and downstream were calculated for WGBS and CAPS data, and 5hmC modification levels were used as the target data of the model. They were classified according to chromosome numbers to construct training and test sets respectively. S2. Feature representation: The DNA sequences of 50 bp upstream and downstream and the methylation level data features of 50 CpG sites upstream and downstream in the training set and the test set are combined for representation to construct the feature input matrix required for training; S3. Model training and tuning: A convolutional neural network model and a mean square error loss function were used for training and parameter tuning to predict the modification level of 5hmC.
2. The method for predicting epigenetic modification signals on a DNA sequence according to claim 1, characterized in that: In the step S1, the 5hmC modification level on CpG is extracted from the CAPS data as the target data, the sites with chromosome numbers 1 and 2 are used as the test set, and the sites with chromosome numbers 3 and 4 are used as the training set.
3. The method for predicting epigenetic modification signals on a DNA sequence according to claim 1, characterized in that: In step S2, the DNA methylation levels of 50 CpGs upstream and downstream of each CpG are extracted from the WGBS data, and the DNA sequences of 50 bp upstream and downstream of each CpG site are extracted, and the DNA sequences are encoded using the One-hot encoding method to form a 4×100 data matrix. The DNA methylation levels extracted from the WGBS are merged to form a 5×100 data matrix, and the formula is as follows: 。 4. The method for predicting epigenetic modification signals on a DNA sequence according to claim 1, characterized in that In step S3, the feature data matrix The convolutional neural network model is input, and the 5hmC modification level is used as the output. The CNN model is trained to learn the relationship between features and 5hmC. The model parameters are adjusted according to the prediction accuracy. The loss function used by the model during training is MSE, which is calculated as: ; Where n is the number of samples, is the actual 5hmC modification level of the i-th sample, The model predicts the 5hmC modification level for the i-th sample.
5. The method for predicting epigenetic modification signals on a DNA sequence according to claim 1, characterized in that In step S3, the trained CNN model is used to predict the test set data to evaluate the model's predictive ability for 5hmC modification levels.
Citation Information
Patent Citations
DNA methylation extension method
CN110060736A
Chromatin topological correlation domain boundary prediction method based on multi-modal fusion
CN115831217A