DNA methylation site recognition and detection system based on data enhancement and Bert model
Through the method of data cleaning, enhancement and dimensionalization combining with the Bert model, data scarcity and interpretability problems in DNA methylation site recognition are solved, and better detection effects are achieved and generalized to RNA methylation site recognition is achieved.
Patent Information
- Application Number
- CN202311536477.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-07-29
AI Technical Summary
There are problems in the field of DNA methylation site recognition, such as scarcity of data, poor interpretability and lack of generalization ability.
Using a combination of data cleaning, data enhancement, data dimensionalization and Bert model, high-quality DNA sequences are obtained through the data cleaning module, the data enhancement module enhances the data volume, and the data dimensionalization module upgrades the data into a large-scale matrix, and uses the Bert model to extract features and combines cross entropy for methylation site recognition.
It improves the detection effect of DNA methylation site recognition and can generalize to the field of RNA methylation site recognition, enhancing the interpretability and detection results of the system.
Smart Images

Figure CN120388614A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of DNA methylation site recognition, and in particular, to a DNA methylation site recognition and detection system based on data augmentation and the Bert model. Background Art
[0002] DNA methylation site recognition is an important task in the field of bioinformatics, which involves determining which sites on the DNA molecule are methylated, and this is of great significance in gene regulation and disease research. Deep learning techniques have made significant progress in this field. CNNs are widely used for DNA methylation site recognition, and they can effectively capture local features in DNA sequences. Researchers have designed various CNN architectures to improve the recognition performance of methylation sites. For example, recurrent neural networks (RNNs): RNNs are very useful for processing sequence dependencies in DNA sequences. Variants of RNNs such as long short-term memory networks (LSTMs) and gated recurrent units (GRUs) have been successfully applied to the task of DNA methylation site recognition. In addition, some researchers have studied the fusion of deep learning models: combining CNNs and RNNs to fully utilize their advantages and improve the model performance. At the same time, applying the attention mechanism helps the model focus on important DNA sequence fragments related to methylation and improve the recognition performance.
[0003] However, there are also some deficiencies in the existing technologies: 1) Data scarcity problem: The acquisition cost of methylated site labeled data is high, so data scarcity is a major challenge. More publicly available datasets are needed in the future to promote the development of models. 2) Interpretability: Deep learning models are usually regarded as black box models and it is difficult to explain their decision-making processes. In biological research, understanding the decisions of the model is crucial for revealing biological knowledge. 3) Generalization ability: Deep learning models need to have good generalization ability to adapt to data under different species and different experimental conditions. In addition, existing machine learning-based methods usually use manually designed features such as k-mer combinations and statistical information, which limits their performance. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a DNA methylation site recognition and detection system based on data augmentation and the Bert model to solve the problems of data sparsity, poor interpretability, and lack of generalization ability in the field of DNA sequence methylation site recognition.
[0005] To achieve the above-mentioned invention purpose, the present invention provides a DNA methylation site recognition and detection system based on data augmentation and the Bert model, and the system includes:
[0006] A data cleaning module, which is used to obtain the original DNA sequence data, perform cleaning operations on the original DNA sequence data to obtain high-quality DNA sequences, and then perform methylation sequencing to obtain DNA methylation sequence data;
[0007] A data augmentation module, which is used to perform dictionary mapping on the DNA methylation sequence data and perform multiple self-replications on the DNA methylation sequence to enhance the DNA methylation sequence;
[0008] A data dimensionality increase module, which is used to increase the dimensionality of the DNA methylation sequence data into a large-scale matrix;
[0009] An encoding module, which is used to extract the feature vectors of the DNA methylation sequence data through the Bert model;
[0010] A feature extraction module, which is used to input the extracted feature vectors into a fully connected network including four layers, and predict the samples with methylation and the samples without methylation through the network;
[0011] A classification module, which is used to quantify the difference between the prediction result of the feature extraction module and the target result;
[0012] A testing module, which is used to input the test data into the trained model to judge the methylation sites of different DNA methylation species.
[0013] Furthermore, the data cleaning module performs cleaning operations on the original DNA sequence, including removing low-quality bases, trimming primer sequences, and correcting sequencing errors.
[0014] Furthermore, the data augmentation module uses a dictionary to represent the mapping between amino acid bases in the DNA methylation sequence data and their corresponding digital indexes, and performs self-replication of the DNA methylation sequence.
[0015] Furthermore, the data dimensionality increase module is specifically used to adopt the adaptive embedding Embedding method to embed the sample vectors of the DNA methylation sequence that has completed dictionary mapping and self-replication into a two-dimensional large-scale matrix.
[0016] Furthermore, the encoding module is specifically used to perform the following steps:
[0017] S1. In the masked language modeling task, randomly select the DNA methylation sequence vector, replace the dictionary data mapped in the sequence with the MASK token, and the goal of the model is to predict the replaced data according to the context of the DNA methylation sequence;
[0018] S2. Create a mask matrix, set the values at the positions filled with MASK tokens to 1, and the values at other positions to 0. During the attention calculation, multiply it element-wise with the attention weight matrix, so that the attention weights at the positions filled with MASK tokens are 0, and exclude them from the attention calculation;
[0019] S3. When calculating the attention mechanism, calculate the scalar product of the query vector Q and the key vector K, then scale the calculation result, and introduce a scaling factor. The attention scores go through a softmax operation to obtain normalized attention weights, and the attention weights are used to perform weighted summation on the value vector V to obtain the final attention representation. At the same time, a multi-head attention mechanism and a fully connected network are added to obtain the encoded data.
[0020] Furthermore, in the feature extraction module, both the initial layer and the final layer are linear transformation layers for performing linear transformation on the data. The second layer uses a Relu layer to accelerate convergence, and the third layer combines a Dropout mechanism to reduce overfitting.
[0021] Furthermore, the classification module is trained using the Adam optimizer, and cross-entropy is used for binary classification tasks to quantify the difference between the predicted result and the target result.
[0022] Compared with the prior art, the beneficial effects of the present invention are:
[0023] A DNA methylation site recognition and detection system based on data augmentation and the Bert model provided by the present invention first obtains high-quality DNA methylation sequence data through a data cleaning module, and enhances it through a data augmentation module to ensure an increase in the amount of data. Then, the DNA methylation sequence data is dimensionally elevated into a large-scale matrix through a data dimensionality elevation module, thereby increasing the receptive field of the data, improving the effectiveness of system feature detection, and improving the interpretability of the system. Then, the encoding module uses the Bert model to extract rich features and combines cross-entropy to identify methylation species sites. By combining data augmentation and the Bert model, the detection result of the system is better, and at the same time, it can be generalized to the field of RNA methylation site recognition, and the detection effect is good. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only the preferred embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0025] Figure 1It is a schematic diagram of the overall structure of a DNA methylation site recognition and detection system based on data augmentation and the Bert model provided by an embodiment of the present invention.
[0026] In the figure, A is the data cleaning module, B is the data augmentation module, C is the data dimensionality increase module, D is the encoding module, E is the feature extraction module, F is the classification module, and H is the testing module. Detailed implementation manners
[0027] The principles and features of the present invention will be described below with reference to the accompanying drawings. The listed embodiments are only used to explain the present invention and are not intended to limit the scope of the present invention.
[0028] Refer to Figure 1 , this embodiment provides a DNA methylation site recognition and detection system based on data augmentation and the Bert model. The system specifically includes a data cleaning module A, a data augmentation module B, a data dimensionality increase module C, an encoding module D, a feature extraction module E, a classification module F, and a testing module H.
[0029] Among them, the data cleaning module A is used to obtain the original DNA sequence data. The original DNA sequence data can be obtained from a laboratory or a public database and includes the sequences of genomes, chromosomes, genes, or other DNA fragments. Perform cleaning operations on the original DNA sequence data, such as removing low-quality bases, trimming appropriate primer sequences, correcting sequencing errors, etc., to obtain high-quality DNA sequences. Then perform specific methylation sequencing, including methods such as BS-seq (Bisulfite Sequencing), RRBS (Reduced Representation Bisulfite Sequencing), or other methylation-sensitive sequencing methods, as well as data alignment methods, to obtain DNA methylation sequences. The DNA methylation sequences include the bases A, C, T, and G.
[0030] The DNA methylation sequences are composed of text strings using the bases A, C, T, and G. Usually, these sequences are converted into digital attributes to adapt to machine learning methods. In deep learning models, DNA methylation sequences are usually used to extract semantic information. However, due to their short length, the extracted semantic information may not be rich enough. The work of the data augmentation module B includes two parts. One is to perform dictionary mapping on the DNA methylation sequence data. Specifically, it can be a mapping between amino acid bases and their corresponding digital indexes using a dictionary. The other is to perform multiple self-replications on the DNA methylation sequences to enhance the DNA methylation sequences.
[0031] The data dimension elevation module C is used to elevate the DNA methylation sequence data to a large-scale matrix. In the data dimension elevation module C, in this embodiment, an adaptive embedding method is adopted, and each of the four nucleotide characters is respectively associated with a vector, which can be specifically realized by adding the transferred randomly initialized vector retrieved from the lookup table to the position of the in-sequence vector. Subsequently, each vector dynamically adjusts its value through backpropagation according to the given task during the training process, and multiple replicated DNA methylation sequences are embedded into a two-dimensional matrix. After embedding, the dimension of each sample is relatively large, and the size of the two-dimensional matrix is several times the size of the original one-dimensional DNA methylation sequence. The purpose of the above operations is to greatly increase the features learned by the model.
[0032] The sequence length of DNA methylation sites is usually 41bp, and the sequence length is relatively short. And the labeled data of specific certain species or specific biological conditions is usually very limited. Deep learning models usually require a large amount of training data to perform well. In this embodiment, self-replication of DNA methylation sequences is used to enhance the sequence data, and at the same time, the Embedding method is used to elevate the data to a large-scale matrix to solve the problem of short data sequences.
[0033] The encoding module D is used to extract the feature vectors of DNA methylation sequence data through the Bert model. Bert is a bidirectional language representation model and is constructed based on the transformer architecture. In this embodiment, a pre-trained Bert model, namely DNABert, is adopted. It consists of three transformer layers, each layer having 768 hidden units and having 8 attention heads per layer. In this embodiment, the model is enhanced through data augmentation and embedding into a large-scale matrix. In the encoding module D, a masked language modeling method similar to that used in the original Bert is adopted. The encoding module D is specifically used to perform the following steps:
[0034] S1. In the masked language modeling task, randomly select DNA methylation sequence vectors, replace the dictionary data mapped in the sequence with MASK tokens, and the goal of the model is to predict the replaced data based on the context of the DNA methylation sequence. Through this pre-training task, the model is forced to predict the correct data in the missing context, thereby learning bidirectional context information. By using masking during pre-training, Bert can consider the left and right context information of each data simultaneously, thus better capturing the dependencies between DNA sequence data. Traditional language models with autoregressive properties, such as the decoder part of recurrent neural networks and Transformers, usually only use the left or right context. However, through the masking task, Bert forces the model to predict the missing context during training, thus alleviating the bias problem in autoregressive models. By masking certain words during pre-training, Bert can access various language patterns, thus better capturing the diversity and complexity of language and enhancing its generalization ability in various downstream tasks.
[0035] S2. Create a masking matrix, set the value at the position filled with MASK tokens to 1, and the values at other positions to 0. During the calculation of attention, multiply the masking matrix element-wise with the attention weight matrix, so that the attention weight at the position filled with MASK tokens is 0, excluding it from the attention calculation, ensuring that the model does not assign attention to MASK tokens and MASK tokens do not affect the calculation of the attention weights of other real tokens, which helps to improve the efficiency and performance of the model.
[0036] S3. When calculating the attention mechanism, calculate the scalar product of the query vector Q and the key vector K, then scale the calculation result to prevent the attention scores from becoming too large. At the same time, introduce a scaling factor (usually the square root of the reciprocal of the dimension of the key vector). The attention scores go through a softmax operation to obtain normalized attention weights, and the attention weights are used to perform a weighted sum on the value vector V to obtain the final attention representation. At the same time, add a multi-head attention mechanism and a fully connected network to obtain the encoded data.
[0037] The feature extraction module E is used to input the extracted feature vectors into a fully connected network consisting of four layers, and predict samples with methylation and samples without methylation through the network. The initial layer and the final layer of the network are both linear transformation models for linearly transforming the data. The second layer uses a Relu layer to accelerate convergence, and the third layer combines a Dropout mechanism to reduce overfitting.
[0038] The classification module F is used to quantify the difference between the prediction result of the feature extraction module and the target result. In this embodiment, the classification module F is trained using the Adam optimizer, which is highly efficient and requires less memory resources, and is very suitable for processing models with large parameters. At the same time, in this embodiment, cross-entropy is used for binary classification tasks to quantify the difference between the prediction result and the target result. This method also includes adjusting the parameters of the model to minimize the difference between the predicted probability and the actual label. This optimization can improve the prediction accuracy of the model for DNA methylation classification.
[0039] The test module H is used to input test data into the trained model and output whether there are methylation sites in different DNA methylation species.
[0040] The system provided in this embodiment first enhances the data by self-replicating the DNA methylation sequence. At the same time, Embedding is used to increase the dimension of the data into a large-scale matrix, converting the original one-dimensional short sequence into a large-scale two-dimensional matrix, which is thousands of times larger than the original one-dimensional short sequence. Analogous to converting a one-dimensional small-scale image into a two-dimensional large-scale image, it increases the receptive field of the data, improves the effectiveness of the system's feature detection, and at the same time improves the interpretability of the detection system. At the same time, the multi-head attention mechanism and the feed-forward fully connected network in the Bert model are used to extract rich features. Cross-entropy is combined for methylation species site recognition. The combination of data enhancement and the Bert model makes the detection result of the system better, and at the same time generalizes to the field of RNA methylation site recognition, and its detection result is very good.
[0041] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A DNA methylation site recognition and detection system based on data augmentation and the Bert model, characterized in that, The system includes: A data cleaning module, which is used to obtain the original DNA sequence data, perform cleaning operations on the original DNA sequence data to obtain high-quality DNA sequences, and then perform methylation sequencing to obtain DNA methylation sequence data; A data enhancement module, which is used to perform dictionary mapping on the DNA methylation sequence data and perform multiple self-replications on the DNA methylation sequence to enhance the DNA methylation sequence; A data dimensionality increase module, which is used to increase the dimensionality of the DNA methylation sequence data into a large-scale matrix; An encoding module, which is used to extract the feature vectors of the DNA methylation sequence data through the Bert model; A feature extraction module, which is used to input the extracted feature vectors into a fully connected network including four layers, and predict the samples with methylation and the samples without methylation through the network; A classification module, which is used to quantify the difference between the prediction result of the feature extraction module and the target result; A testing module, which is used to input the test data into the trained model to judge the methylation sites of different DNA methylation species.
2. The DNA methylation site recognition and detection system based on data augmentation and Bert model according to claim 1, characterized in that The data cleaning module performs cleaning operations on the original DNA sequence, including removing low-quality bases, trimming primer sequences, and correcting sequencing errors.
3. The DNA methylation site recognition and detection system based on data augmentation and Bert model according to claim 1, characterized in that, The data enhancement module uses a dictionary to represent the mapping between amino acid bases in the DNA methylation sequence data and their corresponding digital indexes, and performs self-replication of the DNA methylation sequence.
4. The DNA methylation site recognition and detection system based on data augmentation and Bert model according to claim 3, wherein, The data dimensionality increase module specifically adopts the Adaptive Embedding method to embed the sample vectors of the DNA methylation sequence that has completed dictionary mapping and self-replication into a two-dimensional large-scale matrix.
5. The DNA methylation site recognition and detection system based on data augmentation and Bert model according to claim 3, wherein The encoding module specifically is used to perform the following steps: S1. In the masked language modeling task, randomly select the DNA methylation sequence vector, replace the dictionary data mapped in the sequence with the MASK token, and the goal of the model is to predict the replaced data according to the context of the DNA methylation sequence; S2. Create a masked matrix, set the value at the position filled with the MASK token to 1, and the values at other positions to 0. During the attention calculation, multiply it element by element with the attention weight matrix, so that the attention weight at the position filled with the MASK token is 0 and it is excluded from the attention calculation; S3. When calculating the attention mechanism, calculate the scalar product of the query vector Q and the key vector K, then scale the calculation result, and at the same time introduce a scaling factor. The attention score undergoes a softmax operation to obtain the normalized attention weight, and the attention weight is used for weighted summation of the value vector V to obtain the final attention representation. At the same time, a multi-head attention mechanism and a fully connected network are added to obtain the encoded data.
6. The DNA methylation site recognition and detection system based on data augmentation and Bert model according to claim 1, characterized in that, In the feature extraction module, both the initial layer and the final layer are linear transformation layers, which are used to perform linear transformation on the data. The second layer uses a Relu layer to accelerate convergence, and the third layer combines a Dropout mechanism to reduce overfitting.
7. The DNA methylation site recognition and detection system based on data augmentation and Bert model according to claim 1, wherein The classification module is trained using the Adam optimizer and uses cross-entropy for binary classification tasks to quantify the difference between the prediction result and the target result.