A DNA methylation site prediction method based on secondary fine-tuning and integrated algorithm

By performing fine-grained word segmentation and secondary fine-tuning on DNA promoter sequences, combined with an integrated algorithm, the problem that pre-trained models in existing technologies are not friendly to small data sets is solved, achieving higher DNA methylation site prediction accuracy and generalization ability.

CN119380811BActive Publication Date: 2025-09-09ANHUI AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411482977.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-23
Publication Date
2025-09-09
Estimated Expiration
2044-10-23

AI Technical Summary

Technical Problem

The pre-training models in existing technologies are not accurate enough, especially for small data sets, resulting in insufficient accuracy in predicting DNA methylation sites.

Method used

A method based on secondary fine-tuning and ensemble algorithm was adopted to pre-train and fine-tune the DNA promoter sequence through 1-mer, 3-mer, and 5-mer word segmentation. Combined with the BERT model, the model was used for fine-grained word segmentation and adaptive learning using the dataset of the UCSC database. A soft voting strategy was performed in the ensemble algorithm to obtain the final prediction results.

Benefits of technology

The model's prediction accuracy and generalization ability on small data sets are improved, the risk of overfitting is reduced, and the overall accuracy and performance of DNA methylation site prediction are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119380811B_ABST
    Figure CN119380811B_ABST
Patent Text Reader

Abstract

The present invention discloses a DNA methylation site prediction method based on secondary fine-tuning and an integrated algorithm, belonging to the field of bioinformatics technology. The present invention uses human DNA promoter sequences obtained from the UCSC database, after 1-mer, 3-mer, and 5-mer segmentation, as a corpus to pre-train the BERT model, forming a Promoter-BERT model. This model captures and abstracts deep features in DNA promoter sequences, providing an efficient and robust initial state for subsequent fine-tuning. The Promoter-BERT model is first fine-tuned using the three largest datasets of three methylation types, and then fine-tuned again using 14 datasets with smaller data volumes. This allows the model to focus more on learning the unique features of the dataset, thereby better adapting to specific tasks. This helps the model achieve higher accuracy and performance on the target task and reduces the risk of overfitting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioinformatics, and in particular to a DNA methylation site prediction method based on secondary fine-tuning and integrated algorithms. Background Art

[0002] DNA methylation is a key epigenetic modification that affects gene expression by adding methyl groups to DNA without changing its sequence. It plays a vital role in the development of organisms, significantly affecting the normal regulation of transcription, embryonic development, genomic imprinting, genome stability, and chromatin structure. In addition, changes in DNA methylation patterns affected by environmental factors and aging are associated with the occurrence and progression of diseases, especially cancer. The main types include 4-methylcytosine (4mC), 5-hydroxymethylcytosine (5hmC), and N6-methyladenosine (6mA). Therefore, correctly identifying and predicting DNA methylation sites is of great significance for revealing their functions in biology and disease treatment.

[0003] Over the past few decades, significant progress has been made in the prediction of DNA methylation sites based on traditional machine learning and deep learning methods. For example, for predicting DNA methylation at a single site, some researchers developed the 4mcPred-SVM system, which effectively identifies DNA 4mC sites genome-wide by integrating complex feature representation techniques and a two-stage feature enhancement process. The hybrid CNN-RNN architecture used in DeepDNA4mC and the multi-scale receptive field structure in MSNet-4mC, by enhancing the self-attention mechanism, demonstrate advanced deep learning strategies for pattern representation and sequence relationship perception. DeepTorrent combines multiple feature encoding methods and integrates convolutional neural networks with bidirectional long short-term memory networks, advancing the field. For predicting DNA methylation at multiple sites, other researchers have introduced iDNA-MT, a model that primarily utilizes multi-task learning and bidirectional gated recurrent units (BGRUs) to directly extract shared information from raw DNA sequences of different species. Some researchers have also proposed iDNA-ABF, which uses BERT as a multi-scale biolinguistic learning model to effectively establish a mapping between natural language and biological terms, accurately translating methylation-related sequence determinants into corresponding functional roles.

[0004] Traditional DNA methylation site prediction models rarely use pre-trained models, or even no pre-trained BERT model, which limits the accuracy of the results. Existing methods for increasing prediction model accuracy rely solely on a single fine-tuning step, which works well on large datasets but is less effective on smaller datasets. To address this issue, we propose a DNA methylation site prediction method based on secondary fine-tuning and an ensemble algorithm. Summary of the Invention

[0005] The technical problem to be solved by the present invention is how to solve the problem that the pre-training model in the existing technology is not accurate enough and is not friendly to small data sets, and provides a DNA methylation site prediction method based on secondary fine-tuning and integrated algorithm.

[0006] like Figure 6 As shown, the present invention solves the above technical problems through the following technical solutions, which include the following steps:

[0007] S1: Data Preparation

[0008] DNA promoter sequence data were collected from the UCSC database to form the corpus of the pre-training model, and data on 4mc sites, 5hmc sites, and 6ma sites were collected to form the methylation site prediction dataset;

[0009] S2: Feature segmentation

[0010] Each DNA promoter sequence is segmented into 1-mer, 3-mer, and 5-mer to obtain the promoter base sequence after 1-mer, 3-mer, and 5-mer segmentation;

[0011] S3: Model Pre-training

[0012] The promoter base sequences after 1-mer, 3-mer, and 5-mer segmentation were used as the corpus to input into the BERT model for pre-training to obtain the Promoter-BERT model;

[0013] S4: First fine-tuning

[0014] The largest dataset for each methylation site was selected for the first fine-tuning of the Promoter-BERT model, resulting in the first fine-tuned 1-mer, 3-mer, and 5-mer Promoter-BERT models.

[0015] S5: Second fine-tuning

[0016] For different methylation types, the Promoter-BERT model after the first fine-tuning was iteratively fine-tuned using the remaining smaller datasets to obtain the second fine-tuned Promoter-BERT models for 1-mer, 3-mer, and 5-mer.

[0017] S6: Ensemble Learning

[0018] Using a soft voting strategy, the Promoter-BERT models after the second fine-tuning are integrated into the ensemble algorithm to obtain the final ensemble model.

[0019] S7: Methylation site prediction

[0020] The integrated model is used to predict the methylation sites of the DNA promoter sequence to be predicted to obtain a prediction result.

[0021] Furthermore, in step S2, 1-mer is a single base, 3-mer is three consecutive bases, and 5-mer is five consecutive bases. Each DNA promoter sequence is segmented by 1-mer, 3-mer and 5-mer respectively through three word segmenters to obtain the promoter base sequence of each DNA promoter sequence under the three word segmentation methods.

[0022] Furthermore, in step S4, the specific process of the first fine-tuning is as follows:

[0023] S41: The datasets with the largest data volume were selected from the three methylation datasets of 4mC, 5hmC, and 6mA;

[0024] S42: According to the length of the DNA promoter sequence in the data set in step S41 and the performance limitation of the device, the training cycle is set to obtain the training method parameters of the training process;

[0025] S43: Based on the training method parameters of step S42, the DNA promoter sequence is divided into 1-mer, 3-mer and 5-mer as words for labeling;

[0026] S44: Finally, the DNA promoter sequence after segmentation of the above three sites is input into the pre-trained Promoter-BERT model for fine-tuning.

[0027] Furthermore, in step S5, the specific process of the second fine-tuning is as follows:

[0028] S51: Select initial model

[0029] Use the Promoter-BERT model obtained from the first fine-tuning as the basis;

[0030] S52: Expanding the dataset

[0031] Perform 1-mer, 3-mer, and 5-mer segmentation on the data from the remaining multiple datasets that were not used in the first fine-tuning;

[0032] S53: Targeted fine-tuning

[0033] Fine-tune each dataset separately to adapt to its corresponding features, and obtain the Promoter-BERT model after the second fine-tuning.

[0034] Compared with the prior art, the present invention has the following advantages:

[0035] 1. In the data preparation stage, 1-mer, 3-mer, and 5-mer word segmenters are used for fine-grained word segmentation to retain local features, as well as adaptive learning to enhance learning effects.

[0036] 2. Using the promoter sequences in the UCSC dataset as a corpus to pre-train the domain-specific BERT model enhances the learning and generalization capabilities of the pre-promoter sequence related model.

[0037] 3. The second fine-tuning of the model using a smaller dataset allows the model to focus more on learning the unique features of the dataset, thereby better adapting to specific tasks. This helps the model achieve higher accuracy and performance on the target task and reduces the risk of overfitting. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 1 is an overall framework diagram of the DNA methylation site prediction method based on secondary fine-tuning and integrated algorithm in an embodiment of the present invention;

[0039] Figure 2 1mer, 3mer, and 5mer Promoter-BERT models for the six datasets in the present invention are compared, where a is the 4mC-F.vesca dataset, b is the 5hmc-M.musculus dataset, c is the 6mA-T.thermophile dataset, d is the 6mA-D.melanogaster dataset, e is the 6mA-C.equisetifolia dataset, and f is the 6mA-C.elegans dataset.

[0040] Figure 3 In the embodiment of the present invention, Figure 2 In addition to the 6 datasets, the comparison of the results of the remaining 11 datasets under the Promoter-BERT model of 1mer, 3mer and 5mer is shown in the figure, where ak is the remaining 11 datasets;

[0041] Figure 4: These are the UMAP result images before and after the second fine-tuning on the three data sets in the embodiment of the present invention, wherein a is the UMAP result image of the 4mC-C.equisetifolia data set before the second fine-tuning, b is the UMAP result image of the 4mC-C.equisetifolia data set after the second fine-tuning, c is the UMAP result image of the 4mC-S.cerevisiae data set before the second fine-tuning; d is the UMAP result image of the 4mC-S.cerevisiae data set after the second fine-tuning, e is the UMAP result image of the 6mA-R.chinensis data set before the second fine-tuning, and f is the UMAP result image of the 6mA-R.chinensis data set after the second fine-tuning;

[0042] Figure 5 is the AUROC image of all data sets in the embodiment of the present invention, where aq are the 17 data sets in the embodiment;

[0043] Figure 6 It is a schematic flow chart of the DNA methylation site prediction method based on secondary fine-tuning and integrated algorithm of the present invention. DETAILED DESCRIPTION

[0044] The following is a detailed description of an embodiment of the present invention. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process. However, the protection scope of the present invention is not limited to the following embodiment.

[0045] like Figure 1 As shown, this embodiment provides a technical solution: a DNA methylation site prediction method based on secondary fine-tuning and integrated algorithm, which mainly includes the following steps:

[0046] 1. Data preparation

[0047] DNA promoter sequence data were collected from the UCSC database to form the corpus for the pre-training model. Data on 4mc sites, 5hmc sites, and 6ma sites were collected to form the methylation site prediction dataset.

[0048] 2. Feature segmentation

[0049] When processing DNA promoter sequences, each DNA promoter sequence is segmented into 1-mer, 3-mer, and 5-mer segments, achieving a fine-grained division of the DNA promoter sequence and obtaining the corresponding promoter base sequence. By considering a single base (1-mer), three consecutive bases (3-mer), and five consecutive bases (5-mer), more detailed local features of the DNA promoter sequence can be captured.

[0050] 3. Model Pre-training

[0051] The DNA promoter base sequences after 1-mer, 3-mer, and 5-mer segmentation were used as the corpus and input into the BERT model for pre-training to obtain the Promoter-BERT model.

[0052] 4. First fine-tuning

[0053] The largest dataset for each methylation site was selected for the first fine-tuning of the Promoter-BERT model. The resulting 1-mer, 3-mer, and 5-mer Promoter-BERT models ensured excellent performance on larger datasets while retaining rich prior knowledge to support subsequent fine-tuning.

[0054] 5. Second fine-tuning

[0055] For different methylation types, the Promoter-BERT model after the first fine-tuning is iteratively fine-tuned using the remaining smaller data sets to obtain the 1-mer, 3-mer, and 5-mer Promoter-BERT models after the second fine-tuning to ensure the accuracy of the results.

[0056] 6. Ensemble Learning

[0057] A soft voting strategy is adopted to integrate the Promoter-BERT models of 1-mer, 3-mer, and 5-mer after the second fine-tuning in the ensemble algorithm to obtain the final ensemble model.

[0058] 7. Model evaluation and result analysis: Use indicators such as ACC and AUC to evaluate the integrated model and analyze the results obtained from the above process.

[0059] The specific process of the above steps is as follows:

[0060] Step 1: Data Preparation

[0061] In order to further evaluate the effectiveness of our proposed method compared with the current mainstream methods, we selected the same benchmark dataset introduced by iDNA-ABF. The dataset contains a total of 17 datasets, belonging to three types of DNA methylation. The datasets belonging to 4mc sites include F.vesca, C.equisetifolia, etc.; the datasets belonging to 5hmc sites include M.musculus and H.sapiens; the datasets belonging to 6ma sites include T.thermophile, A.thaliana, etc. In order to pre-train the BERT model, the present invention uses the human promoter dataset from UCSC Genome Browser, which is one of the widely used databases in the field of biology. The specific datasets are shown in Table 1 below.

[0062] Table 1 Specific datasets used (17 in total)

[0063]

[0064]

[0065] Step 2: Feature segmentation

[0066] When processing DNA promoter sequences, each DNA promoter sequence is segmented into 1-mer, 3-mer, and 5-mer segments, achieving a fine-grained division of the DNA promoter sequence. By considering combinations of single bases (1-mer), three consecutive bases (3-mer), and five consecutive bases (5-mer), more detailed local features in the DNA promoter sequence can be captured. Different base combination patterns may correspond to different biological meanings. Through fine-grained tokenization, the model can more flexibly learn and adapt to these complex combination patterns, thereby improving its accuracy in prediction and classification tasks. After fine-grained tokenization, the input sequence is converted into a series of embedded representations with deep semantic information. These embedded representations not only contain the original information of the sequence, but also incorporate the model's understanding and abstraction of the sequence, providing strong support for subsequent analysis and prediction.

[0067] Step 3: Model Pretraining

[0068] In step 3, the base sequences of the three word segments of the DNA promoter sequence are fed into the powerful BERT model for pre-training as a corpus, resulting in the domain-specific pre-trained model Promoter-BERT. This pre-trained model captures and abstracts deep features from DNA promoter sequences, providing an efficient and robust initial state for subsequent fine-tuning.

[0069] Step 4: First fine-tuning

[0070] The present invention selected the data set with the largest amount of data from the three sites of 4mC, 5hmC and 6mA, namely 4mC_F.vesca, 5hmC_M.musculus and 6mA_T.thermophile. Considering that the sequence length is 41 and the performance of our equipment is limited to 32, we divided the training cycles into 4, 6 and 8. For the convenience of recording, we gave them three labels, which we called "training method parameters": 41-32-4, 41-32-6 and 41-32-8. The specific results can be seen in Table 2. Next, we marked the DNA promoter sequence by dividing it into 1-mer, 3-mer and 5-mer as words based on the three parameters in the training method parameters. Then, the "sentences" of these three sites (i.e. the promoter base sequences after the three sites are segmented) were input into the pre-trained Promoter-BERT model for fine-tuning. In order to evaluate the performance of the model, the AUC value was observed, where a higher AUC value indicates a better model effect.

[0071] Table 2 Results of five-fold cross validation and their corresponding 1mer, 3mer, 5mer segmentation results and comprehensive results

[0072]

[0073]

[0074]

[0075]

[0076]

[0077]

[0078] Figure 2 The results of the Promoter-BERT model for 1mer, 3mer, and 5mer on six datasets are shown in Figure 3 The first three datasets are the results of the first fine-tuning on larger datasets, while the last three datasets are the results of the second fine-tuning. As can be seen, there is almost no difference in AUC values ​​for 1mer, 3mer, and 5mer, and the performance is all above 90%, indicating that word segmentation can effectively capture sequence information and make accurate predictions.

[0079] Step 5: Second fine-tuning

[0080] In order to intuitively see the superiority of the present invention, the Promoter-BERT model after the first fine-tuning was fine-tuned again using a dataset with a smaller amount of data, and the features were further visualized.

[0081] We visualized the representation space distribution of the Promoter-BERT model after two fine-tuning on three datasets (4mC_C.equisetifolia, 4mC_S.cerevisiae, and 6mA_R.chinensis), aiming to highlight the difference between the model before and after the two fine-tuning and highlight the importance of the second fine-tuning. We used Uniform Manifold Approximation and Projection (UMAP), a widely used visualization tool that reveals the basic characteristics of data through dimensionality reduction. Figure 4 The attention heatmaps at two scales are visualized (the left side is before the second fine-tuning, and the right side is after the second fine-tuning, Figure 4 In the figure, a and b are 4mC_C.equisetifolia sites, c and d are 4mC_S.cerevisiae sites, and e and f are 6mA_R.chinensis sites). The elements in the heatmap represent the correlation between two positions in the sequence. To this end, we use the attention mechanism to intuitively explain the information learned by our model from 1mer, 3mer to 5mer. The attention heatmaps at two scales are respectively Figure 4 Visualized in . Figure 4 The figure shows the learning information of our model before and after training at three word segmentation scales. We can observe that the scatter plot of the trained model is more dispersed compared to the initial model, which indicates that our model has learned more local discriminative information compared to before training.

[0082] Step 6: Ensemble Learning

[0083] To further improve the prediction performance of the model, we considered combining the prediction results of the three scales after further fine-tuning. Therefore, we chose to use soft voting in the ensemble algorithm to obtain the final prediction results.

[0084] In the first fine-tuning stage, we used the 4mC_F.vesca dataset and set the parameters to 41-32-4. For the 1mer, 3mer, and 5mer models, the AUROC values ​​were 0.9175, 0.9326, and 0.9245, respectively. The AUROC value of the integrated model was 0.9371, which exceeded that of the individual models. When the parameters were adjusted to 41-32-6, the AUROC values ​​of the 1mer, 3mer, and 5mer models were 0.9221, 0.9333, and 0.9223, respectively, and the AUROC value of the integrated model reached 0.9389. Similarly, when the parameters were set to 41-32-8, the AUROC values ​​of the 1mer, 3mer, and 5mer models were 0.9237, 0.9320, and 0.9169, respectively, while the AUROC value of the integrated model was 0.9377, which again exceeded that of the individual models. Similar trends were observed on other datasets. For specific results, see Figure 5 .

[0085] Step 7: Model evaluation and result analysis

[0086] To evaluate the ensemble model, we used several commonly used metrics, including overall accuracy (ACC), Matthews correlation coefficient (MCC), and precision and recall, to evaluate the prediction performance. The formulas for these metrics are as follows:

[0087]

[0088] In the above formula, TP represents the number of true positive examples, FN represents the number of false negative examples, TN represents the number of true negative examples, and FP represents the number of false positive examples. Precision intuitively reflects the ability of the classifier to incorrectly label negative samples as positive samples, while recall reflects the ability to find positive samples. At the same time, we also used the area under the receiver operating characteristic curve (AUROC) (AUROC), which comprehensively evaluates model performance by balancing the true positive rate (also called sensitivity) and the false positive rate. The value of AUROC ranges from 0 to 1, and the closer to 1, the better the model performance. AUROC is robust to class imbalance and data distribution, and is a valuable indicator for comparing the performance of different models. It also helps to determine the optimal classification threshold to balance sensitivity and specificity. The specific AUROC images of 17 data sets are shown below. Figure 5 shown.

[0089] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A DNA methylation site prediction method based on secondary fine-tuning and integrated algorithm, characterized in that: The following steps are involved: S1: Data Preparation DNA promoter sequence data were collected from the UCSC database to form the corpus of the pre-training model, and data on 4mc sites, 5hmc sites, and 6ma sites were collected to form the methylation site prediction dataset; S2: Feature segmentation Each DNA promoter sequence is segmented into 1-mer, 3-mer, and 5-mer to obtain the promoter base sequence after 1-mer, 3-mer, and 5-mer segmentation; S3: Model Pre-training The promoter base sequences after 1-mer, 3-mer, and 5-mer segmentation were used as the corpus to input into the BERT model for pre-training to obtain the Promoter-BERT model; S4: First fine-tuning The largest dataset for each methylation site was selected for the first fine-tuning of the Promoter-BERT model, resulting in the first fine-tuned 1-mer, 3-mer, and 5-mer Promoter-BERT models. S5: Second fine-tuning For different methylation types, the Promoter-BERT model after the first fine-tuning was iteratively fine-tuned using the remaining smaller datasets to obtain the second fine-tuned Promoter-BERT models for 1-mer, 3-mer, and 5-mer. S6: Ensemble Learning Using a soft voting strategy, the Promoter-BERT models after the second fine-tuning are integrated into the ensemble algorithm to obtain the final ensemble model. S7: Methylation site prediction The integrated model is used to predict the methylation sites of the DNA promoter sequence to be predicted to obtain a prediction result.

2. A DNA methylation site prediction method based on secondary fine-tuning and integrated algorithm according to claim 1, characterized in that: In step S2, 1-mer is a single base, 3-mer is three consecutive bases, and 5-mer is five consecutive bases. Each DNA promoter sequence is segmented by 1-mer, 3-mer and 5-mer respectively through three segmentation methods to obtain the promoter base sequence of each DNA promoter sequence under the three segmentation methods.

3. A DNA methylation site prediction method based on secondary fine-tuning and integrated algorithm according to claim 2, characterized in that: In step S4, the specific process of the first fine-tuning is as follows: S41: The datasets with the largest data volume were selected from the three methylation datasets of 4mC, 5hmC, and 6mA; S42: According to the length of the DNA promoter sequence in the data set in step S41 and the performance limitation of the device, the training cycle is set to obtain the training method parameters of the training process; S43: Based on the training method parameters of step S42, the DNA promoter sequence is divided into 1-mer, 3-mer and 5-mer as words for labeling; S44: Finally, the DNA promoter sequence after segmentation of the above three sites is input into the pre-trained Promoter-BERT model for fine-tuning.

4. A DNA methylation site prediction method based on secondary fine-tuning and integrated algorithm according to claim 3, characterized in that: In step S5, the specific process of the second fine-tuning is as follows: S51: Select initial model Use the Promoter-BERT model obtained from the first fine-tuning as the basis; S52: Expanding the dataset Perform 1-mer, 3-mer, and 5-mer segmentation on the data from the remaining multiple datasets that were not used in the first fine-tuning; S53: Targeted fine-tuning Fine-tune each dataset separately to adapt to its corresponding features, and obtain the Promoter-BERT model after the second fine-tuning.

Citation Information

Patent Citations

  • DNA methylation prediction method and system based on BERT framework

    CN113744805A

  • Computational filtering of methylated sequence data for predictive modeling

    US20200327959A1