Enhancer prediction method based on ensemble learning and deep learning

Through the enhanced predictive method based on integrated learning and deep learning, the Blending-KAN model and Stacking-Auto model are used to solve the prediction difficulties caused by relying on the complete combination of multiple epigenetic signals in the prior art, and the effect of accurately predicting and localizing enhancers when the signal data is incomplete is achieved.

CN120220827APending Publication Date: 2025-06-27XIAN UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510268852.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When predicting and localizing enhancers, the prior art relies on a complete combination of multiple epigenetic signals, resulting in the inability to effectively predict enhancers when signal data is incomplete, which affects the wide range of applications.

Method used

Using an enhancer prediction method based on integrated learning and deep learning, the Blending-KAN model and Stacking-Auto model are used to flexibly adopt a variety of epigenetic signal combinations to predict and locate enhancers, and can still work effectively when the signal data is incomplete.

Benefits of technology

It improves the accuracy and robustness of enhancer prediction, and can obtain reliable prediction results without preparing five epigenetic signals, which enhances the flexibility and wide application of the method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220827A_ABST
    Figure CN120220827A_ABST
Patent Text Reader

Abstract

The invention discloses an ensemble learning and deep learning-based enhancer prediction method, which comprises the following steps of: firstly, performing enhancer prediction on multi-dimensional epigenetic signal data by utilizing a Blending-KAN model so as to identify an enhancer region, and on the basis, aiming at the region which is predicted as the enhancer, segmenting the region into subsequences through a sliding window so as to obtain an enhanced region; the method comprises the following steps: firstly, constructing a plurality of sub-sequences, further predicting the probability that each sub-sequence is an enhancer through a Stacking-Auto model, and finally, accurately positioning a complete enhanced sub-region by adopting a dynamic threshold algorithm based on the probability values; according to the multi-model combination method provided by the invention, the accuracy and fine positioning capability of enhancer identification are effectively improved, and the method is expected to be applied to more complex gene regulation network researches.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of bioinformatics, and particularly relates to a method for enhancer prediction based on ensemble learning and deep learning. Background Art

[0002] Precise regulation of gene expression is crucial for maintaining cell function and organism development. During gene expression, transcriptional regulation plays a central role, which involves complex molecular mechanisms to ensure the activation or inhibition of the expression of specific genes at the right time and place. Regulatory elements such as transcription factors (TFs), promoters, and enhancers together constitute this precise regulatory network. Enhancers are a class of special DNA sequences that can regulate gene transcription over long distances. By interacting with the promoter region, they enhance or inhibit gene expression. The presence and activity of enhancers are crucial for cell-type specific expression patterns, and their dynamic changes during development are also closely related to the occurrence and development of various diseases.

[0003] Although the role of enhancers in transcriptional regulation has been widely recognized, how to accurately predict enhancer regions in the genome remains a challenge. Traditional enhancer identification methods mainly rely on experimental techniques such as ChIP-seq, DNase-seq, and ATAC-seq. Although these methods can provide a large amount of regulatory element information, they usually take a long time, are costly, and are complex to operate. With the development of bioinformatics and computational biology, methods based on machine learning and deep learning have become important tools for enhancer prediction. These methods can predict enhancers and their functional characteristics from a large amount of genomic data by integrating various types of data, such as evolutionary conservation, epigenetic marks, DNA sequence motifs, and transcription factor binding sites.

[0004] In recent years, a variety of computational models have been proposed for enhancer identification and prediction of activity. For example, EnhancerFinder integrates DNA sequence motifs, evolutionary patterns, and functional genomic datasets of different cell types using a multi-kernel learning method to improve the ability to identify enhancers. The limitation of EnhancerFinder is that this method can only be used when DNA sequence motifs, evolutionary patterns, and functional genomic datasets are obtained simultaneously. In fact, it is difficult to obtain these data simultaneously. iEnhancer-BERT is based on a pre-trained DNA language model and is fine-tuned on the enhancer identification task through a transfer learning strategy to extract deeper sequence features. iEnhancer-BERT only uses sequences as features to predict enhancers, and the prediction accuracy is 79.3%, which is relatively low. In addition, the iEnhancer-DHF model combines pseudo k-tuple nucleotide composition (PseKNC) and the FastText method for feature extraction and predicts enhancers through a deep neural network (DNN) model. The accuracy of iEnhancer-DHF is relatively low. Using an independent dataset, the accuracy of the iEnhancer-DHF model in the first layer and the second layer is 83.21% and 67.54% respectively.

[0005] DECODE uses DNase-seq and ChIP-seq data of H3K27ac, H3K4me3, H3K4me1, and H3K9ac to train a deep neural network for accurately predicting cell type-specific enhancers. In addition, DECODE also implements a weakly supervised object detection framework for precisely locating enhancer boundaries, thereby improving the annotation resolution. This method combining deep learning and large-scale functional data not only improves the accuracy of enhancer prediction but also provides a new perspective for the functional study of enhancers.

[0006] DECODE has made significant progress in enhancer prediction. However, DECODE depends on the complete combination of the above five epigenetic signals. In reality, due to experimental conditions, cost limitations, or data availability issues, these five signal data are not necessarily complete. At this time, DECODE cannot be used to predict enhancers, which affects the wide application of DECODE. Summary of the Invention

[0007] Aiming at the problems existing in the prior art, the present invention proposes an enhancer prediction method based on ensemble learning and deep learning. This method can flexibly use a variety of epigenetic signal combinations to predict enhancers, and can effectively predict and locate enhancers in the case of incomplete epigenetic signal data, with higher accuracy, greater robustness, and more flexibility than existing methods.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] An enhancer prediction method based on ensemble learning and deep learning, specifically including the following steps:

[0010] Step 1: Download data from the ENCODE data portal, determine positive and negative samples, and process the signal data of the samples.

[0011] Step 2: Construct a Blending-KAN model to predict whether the samples obtained in Step 1 contain enhancers.

[0012] Step 3: For the samples predicted to contain enhancers in Step 2, use the Stacking-Auto model to determine the start and end positions of the enhancers, which is enhancer localization.

[0013] Furthermore, the specific implementation of Step 1 is as follows:

[0014] Step 1.1: Download data: Download various data of human cell lines HCT116 and A549 from the ENCODE data portal. The collected data includes STARR-seq data, chromatin accessibility DNase-seq, and ChIP-seq data of H3K27ac, H3K4me3, H3K4me1, and H3K9ac.

[0015] Step 1.2: Determine positive and negative samples: First, identify the overlapping regions of DNase-seq peaks and STARR-seq peaks. As long as there is an overlap in position, take out the overlapping regions, which are the overlapping regions. If the overlapping region has an intersection with any one of the ChIP-seq peaks of H3K27ac, H3K4me3, H3K4me1, and H3K9ac, it is regarded as an enhancer region, that is, a positive sample. Negative samples are randomly selected from outside the enhancer region. Both positive and negative samples are extended 2000bp to the left and right with the center point of the overlapping region. The length of positive and negative samples is 4000bp, and the quantity ratio of positive and negative samples is 1:10.

[0016] Step 1.3: Process the signal data of the sample data. For each kind of signal data of each sample obtained in Step 1.2, there are 4000 signal values in the 4000bp region. Aggregate every 10bp region and take the average value. The length of each kind of signal data of each sample after aggregation is 400 average values, that is, the calculated feature dimension of each sample is 400.

[0017] Furthermore, the specific implementation of Step 2 is as follows:

[0018] Step 2.1, construct the Blending-KAN model, which is a new classifier based on the AutoGluon framework and the KAN model;

[0019] Step 2.2, dataset division in the first layer. Divide the positive and negative samples obtained from the human cell line HCT116 into two parts according to a ratio of 7:3: The first part is used as the training set of the base classifier, and the second part is used as the test set of the base classifier. At the same time, the prediction results of each base classifier on the test set will be used as the input features of the second-layer meta-classifier;

[0020] Step 2.3, train the base classifiers in the first layer. In the first layer (layer 1), use the AutoGluon module to train multiple models with different characteristics, which are used as components of the base classifier;

[0021] During the process of training the model, to solve the problem of class imbalance, weights are assigned to the positive and negative samples. The weight ratio of positive and negative samples is set to 10:1. Subsequently, the processed data is converted into the TabularDataset format of AutoGluon, and a binary classification model is constructed using TabularPredictor. After training is completed, the model is saved;

[0022] Step 2.4, use TabularPredictor of AutoGluon to predict the test set data. All models trained by AutoGluon are used as base classifiers, and the prediction results of each base classifier on the test set are used as the input of the second-layer meta-classifier;

[0023] Step 2.5, merge the output results of Step 2.4 into a new feature matrix as the input data of the second layer;

[0024] Step 2.6, train the KAN model in the second layer. Randomly divide the new feature matrix merged in Step 2.5 into 5 subsets of similar sizes. In each fold, select one of the subsets as the validation set, and the remaining 4 subsets are merged as the training set. During the training process of each fold, the model is optimized through the PyTorch framework, and the Adam optimizer and cross-entropy loss function are used to train the network. During the training process, the model updates the parameters through the mini-batch gradient descent method;

[0025] Step 2.7, predict the regions containing enhancers in the second layer. In the prediction stage, the model makes inferences on the validation set in the non-gradient update mode, calculates the prediction accuracy, and also evaluates the ROCAUC and PR AUC metrics. The prediction result is whether a DNA region with a length of 4000bp contains an enhancer.

[0026] Further, the specific implementation of step 3 is as follows:

[0027] Step 3.1, establish a benchmark dataset for the Stacking-Auto model. The benchmark dataset introduced in the iEnhancer-2L paper is adopted. This dataset contains enhancer sequences collected from nine different cell lines, and these sequences are divided into 200bp segments. In addition, highly similar DNA sequences are removed using CDHIT. The constructed benchmark dataset is divided into two parts: a training set and an independent test set. The training set contains enhancer sequences and non-enhancer sequences, and the independent test set contains enhancer sequences and non-enhancer sequences.

[0028] Step 3.2, in the task of processing 200bp subsequences, first input the subsequences of the training set into the pre-trained DNABERT-2 model, and output the embedding representation of each 200bp subsequence as the original input feature.

[0029] Step 3.3, select the LightGBM model under the AutoGluon framework as the base learner of the first-layer Stacking model, obtain the prediction results of the first-layer LightGBM model on the training set, and merge these prediction results with the original input features to construct a new feature set. Subsequently, this feature set is input into the second-layer model to generate the final prediction output, obtaining the probability that each 200bp sequence is an enhancer. After training, use the predict method of AutoGluon to predict the independent test set and comprehensively evaluate the model performance through the evaluate method.

[0030] Step 3.4, use the sliding window method to segment the DNA sequence corresponding to the 4000bp region containing enhancers obtained in step 2.7 and extract subsequences.

[0031] Step 3.5, input each subsequence obtained in step 3.4 into the DNABERT2 model for feature extraction to generate the embedding representation characterizing the subsequence characteristics. These embedding representations are then processed by the Stacking-Auto model to generate the probability that each subsequence is an enhancer, obtaining a total of 77 probability values.

[0032] Step 3.6, adopt the dynamic threshold localization algorithm to determine the enhancer boundary. The dynamic threshold Threshold can be expressed by the following formula:

[0033] Threshold = μ + 0.7σ

[0034] where the calculation formulas for the mean μ and the standard deviation σ are:

[0035]

[0036] Among them, P i represents the probability value of the i-th subsequence, and n is the total number of subsequences;

[0037] Based on this screening criterion, adjacent high-probability subsequences are merged to form continuous enhancer regions. For each merged enhancer region, its average probability is calculated. Among all enhancer regions, the enhancer region with the highest average probability is selected as the positioning target of the enhancer. Finally, according to the relative position of the selected region in the original 4000bp sequence, the start and end coordinates in the genome are accurately calculated, and the start and end positions of the enhancer can be determined. The calculation formulas for the start and end coordinates are as follows:

[0038] Start Coordinate=Original Start Position+start×20

[0039] End Coordinate=Original Start Position+(end+1)×200-1

[0040] Among them, Original Start Position is the start coordinate of the 4000bp sequence in the genome, and start and end are the start and end subsequence indices of the merged enhancer region, respectively.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] (1) The present invention proposes a new classifier based on the AutoGluon framework and the KAN model. It integrates multiple base classifiers through model fusion technology, significantly improving the accuracy and robustness of enhancer prediction, and reliable prediction results can be obtained without having to prepare five epigenetic signals.

[0043] (2) The present invention designs a Stacking-Auto sequence classifier, which combines DNABERT-2 feature extraction and the Stacking method under the AutoGluon framework. This design improves the generalization performance of the model through the collaborative work of two-layer learners, optimizes the model training and prediction processes, and realizes the accurate positioning of potential enhancers.

[0044] (3) The present invention proposes a method based on enhancer probability values and a dynamic threshold algorithm for precisely locating the specific positions of enhancers. This method first segments the sequence data through a sliding window technique. Subsequently, it calculates the probability that each subsequence is an enhancer in combination with the Stacking-Auto model. Finally, it accurately identifies the active enhancer regions through the dynamic threshold algorithm. This method provides precise genomic annotation for subsequent functional research and genetic variation analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a flowchart of an enhancer prediction method based on ensemble learning and deep learning according to the present invention;

[0046] Figure 2 is a structural diagram of the Blending-KAN model in an enhancer prediction method based on ensemble learning and deep learning according to the present invention;

[0047] Figure 3 is a process framework diagram of the Stacking-Auto model in an enhancer prediction method based on ensemble learning and deep learning according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0049] As Figure 1 shown, an enhancer prediction method based on ensemble learning and deep learning specifically includes the following steps:

[0050] Step 1, Data set preprocessing

[0051] Step 1.1, Download data: Download various data of human cell lines HCT116 and A549 from the ENCODE data portal (https: / / www.encodeproject.org / ). The collected data includes STARR-seq data, chromatin accessibility DNase-seq, ChIP-seq data of H3K27ac, H3K4me3, H3K4me1, and H3K9ac;

[0052] Step 1.2, Determine positive and negative samples: First, identify the overlapping regions of DNase-seq peaks and STARR-seq peaks. As long as there is an overlap in position, extract the overlapping region, which is the overlapping region. If this overlapping region intersects with any one of the ChIP-seq peaks of H3K27ac, H3K4me3, H3K4me1, and H3K9ac, it is regarded as an enhancer region, that is, a positive sample. Negative samples are randomly selected from outside the enhancer region. Both positive and negative samples are extended 2000bp to the left and right from the center point of the overlapping region. The length of positive and negative samples is 4000bp, and the quantity ratio of positive and negative samples is 1:10. Finally, five types of signal data of 8350 positive samples and 83500 negative samples are obtained on the human cell line HCT116 for model construction;

[0053] Step 1.3, Process the signal data of the sample data. For each type of signal data of each sample obtained in Step 1.2, in the 4000bp region, there are 4000 signal values. Aggregate every 10bp region and take the average value. After aggregation, the length of each type of signal data of each sample is 400 average values;

[0054] Step 2, As Figure 2 shown, construct the Blending-KAN model, flexibly adopt multiple signal combinations to predict enhancers, so as to identify enhancer regions;

[0055] Design the Blending-KAN model to predict whether a 4000bp region contains an enhancer. Sequentially divide the data extracted from k (0 < k <= 5, the Blending-KAN model supports the free combination of five types of signal data) epigenetic signals into a training set and a test set. Use the first part of the training set to train multiple different base classifiers through the AutoGluon module, use these base classifiers to predict the second part of the test set, and generate a new prediction result set. These prediction results will be used as input features for the meta-classifier.

[0056] Step 2.1, Construct the Blending-KAN model, which is a new type of classifier based on the AutoGluon framework and the KAN model. This model uses the idea of integrated learning Blending. It integrates multiple base classifiers through model fusion technology and uses the KAN model as the meta-classifier.

[0057] Step 2.2, Dataset division in the first layer: Divide the positive and negative samples obtained from the human cell line HCT116 into two parts according to a ratio of 7:3. The first part is used as the training set for the base classifier, and the second part is used as the test set for the base classifier. Meanwhile, the prediction results of each base classifier on the test set will be used as the input features for the meta-classifier in the second layer.

[0058] The basis for the 7:3 dataset division ratio is as follows:

[0059] In the Blending algorithm, dividing the dataset is a crucial step because the performance of Blending depends to a large extent on how the training data is divided. To determine the most suitable division ratio for this task, all five types of signal data were used, and five division ratios (5:5, 6:4, 7:3, 8:2, 9:1) were tried, and experiments were conducted on the above model. Meanwhile, the same dataset was used for training on DECODE. As a result, under the 7:3 division ratio, the Blending-KAN model showed the best performance.

[0060] Step 2.3, Training the base classifier in the first layer: In the first layer (layer 1), use the AutoGluon module to train multiple models with different characteristics, which will be used as components of the base classifier.

[0061] During the process of training the model, to address the class imbalance problem, weights were assigned to the positive and negative samples. The weight ratio of the positive and negative samples was set to 10:1. Subsequently, the processed data was converted into the TabularDataset format of AutoGluon, and a binary classification model was constructed using TabularPredictor. 5-fold bagging (num_bag_folds = 5) was adopted to enhance the robustness of the model, and the preset parameter "best_quality" was used to optimize the model performance. After training, the model was saved.

[0062] Step 2.4, Prediction results in the first layer: Use the TabularPredictor of AutoGluon to predict the test set data. All the models trained by AutoGluon are used as base classifiers, and the prediction results of each base classifier on the test set are used as the input features for the meta-classifier. Specifically, the prediction results of each base classifier are stacked by column, and finally a new feature set is formed. This feature set contains prediction information from different base classifiers and is used as the input for the meta-classifier in the second layer (layer 2).

[0063] Step 2.5, the output result of Step 2.4 is used as the input data of the second layer. Each type of signal data is predicted by an independent classifier set. This model supports flexibly adopting various signal combinations to accurately predict enhancers. Finally, the feature sets of all selected k (0 < k <= 5) signal data are merged into a new feature matrix, which is used to construct the input of the second layer (layer2).

[0064] Step 2.6, train the KAN model in the second layer. In the training stage, the meta-classifier adopts the KAN network. KAN is a new type of deep learning model. Compared with the traditional MLP, the core feature of the KAN network is that it places the activation function at the edge of the network (i.e., multiplies with the weights). And these activation functions are learnable and are usually parameterized using B-splines. This design makes the KAN network have higher accuracy and interpretability in dealing with complex function fitting, partial differential equation solving, etc.

[0065] To ensure the reliability of model evaluation, a five-fold cross-validation strategy is adopted. The new feature matrix merged in Step 2.5 is randomly divided into 5 subsets of similar sizes. In each fold, one of the subsets is selected as the validation set, and the remaining 4 subsets are merged as the training set. The StratifiedKFold method is used to ensure that the class label distribution in each fold is consistent with the original data, thus avoiding the impact of data imbalance on the model training results. During the training process of each fold, the model is optimized through the PyTorch framework, and the Adam optimizer and cross-entropy loss function are used to train the network. During the training process, the model updates the parameters through the mini-batch gradient descent method.

[0066] Step 2.7, predict the regions containing enhancers in the second layer. In the prediction stage, the model makes inferences on the validation set in the no-gradient update mode and calculates the prediction accuracy. At the same time, the ROCAUC and PR AUC metrics are also evaluated to comprehensively reflect the classification performance of the model and ensure the accuracy and robustness of the evaluation results; the prediction result is whether a 4000bp-long DNA region contains an enhancer.

[0067] Step 3, enhancer localization

[0068] For the 4000bp-long DNA region predicted to contain an enhancer in Step 2, use the Stacking-Auto model to determine the start and end positions (boundaries) of the enhancer, which is the enhancer localization.

[0069] To localize the boundaries of the enhancer from the regions containing the enhancer, the present invention designs a Stacking-Auto model, as Figure 3As shown below. First, design a Stacking-Auto model to predict the probability that a 200bp subsequence is an enhancer. Second, extract the DNA sequence data of a 4000bp-long DNA region predicted to contain an enhancer. Then, split the DNA sequence data into 200bp subsequences in a sliding window manner (with a step size of 50), and the Stacking-Auto model predicts the probability. Finally, design a dynamic threshold algorithm to determine the boundaries of active enhancers based on the probability.

[0070] The specific approach is as follows:

[0071] Step 3.1, establish a benchmark dataset for the Stacking-Auto model. To train a suitable model to predict whether a 200bp sequence is an enhancer, the benchmark dataset introduced in the iEnhancer-2L paper is adopted. This dataset contains enhancer sequences collected from nine different cell lines, and these sequences are divided into 200bp segments. In addition, CDHIT is used to remove highly similar DNA sequences (similarity > 20%). Finally, the constructed benchmark dataset is divided into two parts: a training set and an independent test set. The training set contains 1484 enhancer sequences, with 742 weak enhancers and 742 strong enhancers each, and also contains 1484 non-enhancer sequences. The independent test set contains 200 enhancer sequences, with 100 weak enhancers and 100 strong enhancers each, and 200 non-enhancer sequences. In subsequent analyses, both strong and weak enhancers are classified as the positive class for unified classification processing;

[0072] Step 3.2, in the task of processing 200bp subsequences, first input the training set subsequences into the pre-trained DNABERT-2 model, and output the embedding representation of each 200bp subsequence as the original input feature for downstream model training.

[0073] Step 3.3, select the LightGBM model under the AutoGluon framework as the base learner of the first-layer Stacking model, obtain the prediction results of the first-layer LightGBM model on the training set, and merge these prediction results with the original input features to construct a new feature set. Subsequently, this feature set is input into the second-layer model to generate the final prediction output, obtaining the probability that each 200bp sequence is an enhancer.

[0074] In the training of the first-layer model, a 10-fold cross-validation method is adopted to ensure the robustness and generalization ability of model training. After the prediction results generated by cross-validation are concatenated with the original features, it provides a richer information representation for the second-layer model. The second-layer model is also based on the AutoGluon framework, and the model with the best performance is selected as the meta-classifier. In the model training stage, the TabularPredictor class of AutoGluon is used to construct the model, and the model is trained through the fit method. After training, the predict method is used to predict the independent test set, and the evaluate method is used to comprehensively evaluate the model performance. The final result of this step is: a Stacking-Auto model is constructed, which can predict the probability that a sequence with a length of 200bp is an enhancer;

[0075] Step 3.4: Use the sliding window method to segment the DNA sequence corresponding to the 4000bp region containing enhancers obtained in Step 2.7, extract subsequences. The length of the sliding window is 200bp, and the step size is 50bp. The 4000bp DNA sequence is divided into 77 200bp subsequences, and there is an overlapping region of 50bp between these subsequences, reducing the possibility of missing potential enhancer regions. In the subsequent steps, the 200bp subsequences will be used as independent samples for feature extraction and obtaining the probability of being an enhancer;

[0076] Step 3.5: Input each subsequence obtained in Step 3.4 into the DNABERT2 model for feature extraction to generate embedding representations characterizing the subsequence characteristics. These embedding representations are processed by the Stacking-Auto model to generate the probability that each subsequence is an enhancer, and a total of 77 probability values are obtained;

[0077] Step 3.6: Use the dynamic threshold localization algorithm to determine the enhancer boundary. The dynamic threshold Threshold can be expressed by the following formula:

[0078] Threshold = μ + 0.7σ (1)

[0079] Among them, the calculation formulas for the mean μ and the standard deviation σ are:

[0080]

[0081] Among them, P i represents the probability value of the i-th subsequence, and n is the total number of subsequences (77 in this embodiment). This dynamic threshold setting can effectively capture relatively high-probability subsequence segments and ensure the identification of regions with high enhancer potential. Its advantage is that it can adapt to the probability distribution differences between different samples, thereby improving the flexibility and accuracy of screening.

[0082] Based on this screening criterion, adjacent high-probability subsequences are merged to form continuous enhancer regions. For each merged enhancer region, its average probability is calculated. Among all enhancer regions, the enhancer region with the highest average probability is selected as the positioning target of the enhancer. Finally, according to the relative position of the selected region in the original 4000bp sequence, the start and end coordinates in the genome are accurately calculated to determine the start and end positions (boundaries) of the enhancer. The calculation formulas for the start and end coordinates are as follows:

[0083] Start Coordinate=Original Start Position+start×20 (4)

[0084] End Coordinate=Original Start Position+(end+1)×200-1 (5)

[0085] Among them, Original Start Position is the start coordinate of the 4000bp sequence in the genome, and start and end are the start and end subsequence indexes of the merged enhancer region respectively.

[0086] The accuracy of the Blending-KAN model is greater than or equal to 94.15% under different combinations of epigenetic signals. Especially when five signals are combined, the accuracy obtained through five-fold cross-validation reaches 99.69%. In cross-cell line prediction, the accuracy of Blending-KAN is greater than or equal to 93.72%. In addition, in the case of adding Gaussian noise, the Blending-KAN model can still maintain an accuracy of 98.74% and above and an AUROC of 99.54% and above. It is particularly stable at low noise levels, showing strong anti-noise ability. When predicting enhancer subsequences, the accuracy, specificity, and MCC of the Stacking-Auto model are 80.50%, 80.50%, and 0.61 respectively.

[0087] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for predicting enhancers based on ensemble learning and deep learning, characterized in that: The specific steps include: Step 1: Download data from the ENCODE data portal, identify positive and negative samples, and process the signal data of the samples; Step 2: construct a Blending-KAN model to predict whether the sample obtained in step 1 contains enhancers; Step 3: For the samples predicted to contain enhancers in step 2, the Stacking-Auto model is used to determine the start and end positions of the enhancers, i.e., enhancer localization.

2. The enhancer prediction method based on ensemble learning and deep learning according to claim 1, characterized in that: The specific steps of step 1 are: Step 1.1, download data: Download various data of human cell lines HCT116 and A549 from the ENCODE data portal. The collected data include STARR-seq data, chromatin accessibility DNase-seq, and ChIP-seq data of H3K27ac, H3K4me3, H3K4me1, and H3K9ac; Step 1.2, determine the positive and negative samples: first identify the overlapping regions of the DNase-seq peak and the STARR-seq peak. As long as there is overlap in position, take out the overlapping region, which is the overlapping region; if the overlapping region intersects with any of the ChIP-seq peaks of H3K27ac, H3K4me3, H3K4me1 and H3K9ac, it is regarded as an enhancer region, that is, a positive sample; negative samples are randomly selected from outside the enhancer region; positive and negative samples are extended 2000bp to the left and right from the center point of the overlapping region, the length of positive and negative samples is 4000bp, and the number ratio of positive and negative samples is 1:10; Step 1.3, process the signal data of the sample data. For each type of signal data of each sample obtained in step 1.2, there are 4000 signal values ​​in the 4000bp region. Aggregate each 10bp region and take the average value. The length of each type of signal data of each sample after aggregation is 400 average values, that is, the calculated feature dimension of each sample is 400.

3. The enhancer prediction method based on ensemble learning and deep learning according to claim 1, characterized in that: The specific steps for step 2 are: Step 2.1, construct the Blending-KAN model. The Blending-KAN model is a new classifier based on the AutoGluon framework and the KAN model. Step 2.2, data set division in the first layer: the positive and negative samples obtained on the human cell line HCT116 are divided into two parts in a ratio of 7:3: the first part is used as the training set of the base classifier, and the second part is used as the test set of the base classifier. At the same time, the prediction results of each base classifier on the test set will be used as the input features of the second-layer meta-classifier; Step 2.3, training the base classifier in the first layer. In the first layer, layer 1, use the AutoGluon module to train multiple models with different characteristics as components of the base classifier. In the process of training the model, in order to solve the problem of class imbalance, weights were assigned to positive and negative samples, and the weight ratio of positive and negative samples was set to 10:

1. Then, the processed data was converted into the TabularDataset format of AutoGluon, and the TabularPredictor was used to build a binary classification model. After the training was completed, the model was saved; Step 2.4, use AutoGluon's TabularPredictor to predict the test set data, use all models trained by AutoGluon as base classifiers, and use the prediction results of each base classifier on the test set as the input of the second-layer meta-classifier; Step 2.5: The output results of step 2.4 are merged into a new feature matrix as the input data of the second layer; Step 2.6, train the KAN model in the second layer, randomly divide the new feature matrix merged in step 2.5 into 5 subsets of similar size, select one subset as the validation set in each fold, and combine the remaining 4 subsets as the training set. During the training process of each fold, the model is optimized through the PyTorch framework, and the Adam optimizer and cross entropy loss function are used to train the network. During the training process, the model updates the parameters through the mini-batch gradient descent method; Step 2.7, predict the region containing enhancers in the second layer. In the prediction stage, the model infers the validation set in the gradient-free update mode and calculates the prediction accuracy. It also evaluates the ROCAUC and PR AUC indicators. The prediction result is: whether a DNA region of 4000 bp in length contains enhancers.

4. The enhancer prediction method based on ensemble learning and deep learning according to claim 1, characterized in that: The specific steps for step 3 are: Step 3.1, build a benchmark dataset for the Stacking-Auto model. The benchmark dataset introduced in the iEnhancer-2L paper was used. This dataset contains enhancer sequences collected from nine different cell lines, and these sequences are divided into 200bp fragments. In addition, CDHIT is used to remove highly similar DNA sequences. The constructed benchmark dataset is divided into two parts: a training set and an independent test set; the training set contains enhancer sequences and non-enhancer sequences, and the independent test set contains enhancer sequences and non-enhancer sequences; Step 3.2, in the task of processing 200bp subsequences, first input the subsequences of the training set into the pre-trained DNABERT-2 model, and output the embedded representation of each 200bp subsequence as the original input feature; Step 3.3, select the LightGBM model under the AutoGluon framework as the basic learner of the first-layer Stacking model, obtain the prediction results of the first-layer LightGBM model on the training set, and merge these prediction results with the original input features to construct a new feature set. Subsequently, this feature set is input into the second-layer model to generate the final prediction output, and the probability of each 200bp sequence being an enhancer is obtained. After the training is completed, use the predict method of AutoGluon to predict the independent test set, and use the evaluate method to comprehensively evaluate the model performance; Step 3.4, using a sliding window method to segment the DNA sequence corresponding to the 4000 bp region containing the enhancer obtained in step 2.7, and extract subsequences; Step 3.5: Input each subsequence obtained in step 3.4 into the DNABERT2 model for feature extraction to generate an embedding representation that characterizes the characteristics of the subsequence. These embedding representations are then processed by the Stacking-Auto model to generate the probability that each subsequence is an enhancer, and a total of 77 probability values ​​are obtained; Step 3.6, a dynamic threshold positioning algorithm is used to determine the enhancer boundary. The dynamic threshold Threshold can be expressed by the following formula: Threshold=μ+0.7σ The calculation formulas for the mean μ and standard deviation σ are: Among them, P i represents the probability value of the i-th subsequence, and n is the total number of subsequences; Based on this screening criterion, adjacent high-probability subsequences are merged to form continuous enhancer regions. For each merged enhancer region, its average probability is calculated. Among all enhancer regions, the enhancer region with the highest average probability is selected as the enhancer positioning target. Finally, according to the relative position of the selected region in the original 4000bp sequence, its start and end coordinates in the genome are accurately calculated to determine the start and end positions of the enhancer. The calculation formula for the start and end coordinates is: Start Coordinate=Original Start Position+start×20 End Coordinate=Original Start Position+(end+1)×200-1 Among them, Original Start Position is the starting coordinate of the 4000bp sequence in the genome, and start and end are the starting and ending subsequence indexes of the merged enhancer region, respectively.

Citation Information

Cited By

  • Corn gene function evaluation method, system and equipment and storage medium

    CN121281648A