Method for predicting interaction between enhancers and promoters based on multi-omics data

Through the deep neural network model, especially the KansformerEPI model, combined with the multiomics data set for preprocessing and training, the problem of the existing technology in inadequate consideration of nonlinearity and not global analysis from the perspective of multicellular lines when integrating features is solved, and accurate prediction of enhancer promoter interactions is achieved.

CN120072031AActive Publication Date: 2025-05-30NORTHEAST FORESTRY UNIV

Patent Information

Application Number
CN202510020342.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-30
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

The prior art does not consider nonlinearity when integrating features, and does not conduct global analysis from the perspective of multicellular lines, making it difficult to accurately identify enhancer promoter interactions.

Method used

Deep neural network model, especially the KansformerEPI model, is adopted, which combines CNN, BiLSTM and Kansformer encoder, and is preprocessed and trained through multi-omics datasets to predict enhancer promoter interactions.

Benefits of technology

It improves the accuracy of the prediction of enhancer promoter interactions, can achieve excellent predictive performance on multiple cell lines, and perform better results on multiple evaluation indicators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072031A_ABST
    Figure CN120072031A_ABST
Patent Text Reader

Abstract

The invention provides a method for predicting interaction of EPIs (Enhancer Promoter) based on multi-omics data, belongs to the technical field of bioinformatics, and solves the problems that nonlinearity is not considered thoroughly during feature integration and global analysis of EPIs is not carried out at the perspective of a multi-cell line in the common technology, and the method comprises the following steps: step 1, constructing a deep neural network model; 2, acquiring multi-omics data and constructing a multi-omics data set; step 3, preprocessing the multi-omics data set; 4, training a deep neural network model based on the preprocessed multi-omics data set; and step 5, based on the trained deep neural network model, performing enhancement promoter interaction prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for predicting enhancer-promoter interactions based on multi-omics data, and belongs to the technical field of bioinformatics. Background Art

[0002] Enhancers and promoters are two important DNA sequence elements in cells. Enhancers play a crucial role in the transcriptional regulation of genes and can enhance the expression of target genes. The promoter, as the "switch" of gene expression, is usually defined as the region from 1500 bp upstream to 500 bp downstream of the transcription start site (TSS). As a cis-regulatory element in the genome, enhancers regulate gene expression through interaction with the promoter of the target gene. Recent studies on three-dimensional genomes have revealed that distal enhancers can interact with proximal promoters to regulate the expression of target genes. Accurately identifying EPIs is crucial for understanding gene regulation, cell differentiation, and disease mechanisms. There are some variations related to cancer susceptibility in enhancer regions. For example, the 8q24 region contains multiple functional enhancers related to various tumor types. However, since enhancers are usually far from their target promoters, accurately identifying EPIs remains challenging.

[0003] Previous studies inferred EPIs through expression quantitative trait loci (eQTLs), but eQTL mapping usually requires a large number of samples and mainly identifies short-distance EPIs. With the development of high-throughput sequencing technologies, researchers can now use techniques such as paired-end tag sequencing (ChIA-PET) and high-throughput chromosome conformation capture (Hi-C) to detect long-distance chromatin interactions, and these techniques can be used to identify EPIs. For example, the application of Hi-C technology has revealed multiple topologically associated domains (TADs). Studies have shown that within these TADs, the interaction frequency of mammalian chromatin is significantly higher than that outside the TADs. However, although these high-throughput technologies have demonstrated powerful capabilities in detecting EPIs, they consume a large amount of time, and the costs of sample preparation, library construction, and sequencing equipment are relatively high.

[0004] To address these problems, many computational methods have been proposed for large-scale, rapid, and accurate identification of EPIs. The data types applied in these methods can be roughly divided into two categories: DNA sequences and epigenetic data. Most of the above-mentioned computational methods are used to predict EPIs constructed with a single cell line, and can only identify the EPIs specific within this cell line. When identifying EPIs in other tissues, the model needs to be retrained. Therefore, it is very necessary to construct a global prediction model for EPIs, that is, a global analysis of EPIs from the perspective of multiple cell lines. Summary of the Invention

[0005] To solve the problems that the common technologies do not consider non - linearity well when integrating features and do not conduct a global analysis of EPIs from the perspective of multiple cell lines, the present invention proposes a method for predicting enhancer - promoter interactions based on multi - omics data.

[0006] The technical solutions adopted by the present invention to solve the above problems are as follows: The present invention specifically includes:

[0007] Step 1: Construct a deep neural network model;

[0008] Step 2: Obtain multi - omics data and construct a multi - omics dataset;

[0009] Step 3: Pre - process the multi - omics dataset;

[0010] Step 4: Train the deep neural network model based on the pre - processed multi - omics dataset;

[0011] Step 5: Predict enhancer - promoter interactions based on the trained deep neural network model.

[0012] Preferably, in Step 1, the deep neural network model is composed of a CNN network, a BiLSTM network, and a Kansformer encoder;

[0013] The CNN network is a convolutional neural network, which is used to identify local patterns and key signal features in the input multi - omics dataset;

[0014] The BiLSTM network is a bidirectional long short - term memory network, which is used to process features in the forward and reverse directions of the multi - omics dataset to enhance the ability to capture time - dependent relationships;

[0015] The Kansformer encoder is used to capture long - range dependencies in the multi - omics dataset.

[0016] Preferably, the Kansformer encoder is composed of the fusion of KAN layers and Transformer, including the alternating stacking of multi - head self - attention layers and KAN layers. The multi - head self - attention layers are used to capture the relationships between features, and the KAN layers are used for predicting enhancer - promoter interaction relationships.

[0017] Preferably, in Step 2, the multi - omics data includes CTCF binding sites, DNase - I signals, 5 histone modification features, and DNA methylation information, and the multi - omics dataset includes Hi - C datasets, CHIA - PET datasets, and feature signals.

[0018] Preferably, Step 3 specifically includes:

[0019] Step 3.1: For the Hi-C dataset and CHIA-PET dataset, convert enhancer-gene pairs into enhancer-promoter pairs. According to the negative sample generation strategy, pair the measured enhancers with genes that have not been experimentally verified to interact with them. The gene selection criterion is that the distance between the gene and the corresponding enhancer is within 95% of the enhancer-gene distance in the positive samples, and complete the preprocessing of the Hi-C dataset and CHIA-PET dataset;

[0020] Step 3.2: For the feature signals, store the feature signals in bigwig files and narrowPeak files, convert the bigwig files and narrowPeak files into pt files storing tensors, and normalize the features to complete the preprocessing of the feature signals.

[0021] Preferably, step 4 specifically includes:

[0022] Step 4.1: Use the preprocessed multi-omics datasets as input sequences, and extract features from the input sequences through a CNN network;

[0023] Step 4.2: The BiLSTM network receives the extracted features and enhances the ability of the extracted features to capture temporal dependencies through forward and backward LSTM networks, and outputs a matrix Z, Z ∈ R l×h which is a matrix containing enhancer and promoter site features, where R is the real number field, l is the number of rows of the matrix, corresponding to the number of sequences, and h is the number of columns of the matrix, corresponding to the dimension of the features;

[0024] Step 4.3: Input the output matrix Z of the BiLSTM network into the Kansformer encoder, and capture the relationships between the features in the output matrix Z through the multi-head self-attention layer to obtain enhanced features

[0025] Step 4.4: Input the enhanced features into the KAN layer for feature integration to obtain the output H of the Kansformer encoder 0 ∈ R l×h ;

[0026] Step 4.5: Use the self-attention embedding method to perform low-dimensional representation learning on H 0 to obtain the aggregated feature H ∈ R 4h ;

[0027] Step 4.6: Input the aggregated feature H into the KAN classifier, calculate the probability p for predicting enhancer-promoter interaction, and predict the distance d from the enhancer to the promoter p ;

[0028] Step 4.7: Based on the calculated enhancer-promoter interaction probability p and the distance d between the enhancer and the promoter p Construct a binary cross-entropy loss function for predicting EPIs and a mean squared error loss function for predicting the enhancer-promoter distance Based on the binary cross-entropy loss function and the mean squared error loss function Construct the total loss function

[0029] Step 4.8: According to the total loss function Train the deep neural network model to obtain the trained deep neural network model;

[0030] The calculation formula for the enhancer-promoter interaction probability p is:

[0031] p = σ(KAN 1 (H)) * (1);

[0032] In formula (1), σ is the sigmoid function, p is the probability of enhancer-promoter interaction, and the range of p values is 0 to 1;

[0033] The distance d between the enhancer and the promoter p The calculation formula is:

[0034] d p = KAN 2 (H)(2);

[0035] The binary cross-entropy loss function The expression is:

[0036]

[0037] In formula (3), y i is the true label (0 or 1) of the i-th sample, and p i is the predicted probability;

[0038] The mean squared error loss function The expression is:

[0039]

[0040] In formula (4), d p is the predicted enhancer-promoter distance, and d t is the true distance;

[0041] The total loss function The expression is:

[0042]

[0043] In formula (5), is the total loss function.

[0044] Preferably, step 4.3 specifically includes:

[0045] Step 4.3.1: Calculate the query Q, key Q, and value vector V based on the matrix Z and the weight matrix for mapping changes;

[0046] Step 4.3.2: Calculate the global relationship matrix M based on the query Q and key Q;

[0047] Step 4.3.3: Enhance the features of the value vector V through the global relationship matrix M to obtain the enhanced features

[0048] The calculation formulas for the query Q, key Q, and value vector V are:

[0049] Q = ZW q , K = ZW k , V = ZW v (6);

[0050] In formula (6), W q , W k and W v are respectively three weight matrices for mapping changes, and K, Q, V are three feature encodings after the linear transformation of the matrix Z;

[0051] The calculation formula for the global relationship matrix M is:

[0052]

[0053] In formula (7), d is the feature dimension of the matrix Z;

[0054] The enhanced features The calculation formula is:

[0055]

[0056] In formula (8), b is the bias.

[0057] Preferably, in step 4.4, the KAN layer is a K-layer KAN network, and a K-layer KAN network includes the nesting of multiple KAN units. The steps of feature integration include:

[0058] Step 4.4.1: For each KAN layer, transform the input through the activation function to obtain the post-activation value

[0059] Step 4.4.2: Sum all the post-activation values of the inputs to obtain the activation value of the neuron and the output H of the Kansformer encoder 0 ∈R l×h ;

[0060] The expression of the KAN network is:

[0061] KAN(Z) = (Φ K-1 , Φ K-2 , …, Φ 1 , Φ 0 )(Z)(9);

[0062] In formula (9), Φ i is the i-th layer of the KAN network. For each KAN layer, the input dimension is n in , and the output dimension is n out . Each layer of the KAN network is defined as:

[0063]

[0064] The formula for calculating the post-activation value is:

[0065]

[0066] Preferably, step 4.5 specifically includes:

[0067] Step 4.5.1: Input H 0 into a two-layer fully connected network to obtain matrix A s ;

[0068] Step 4.5.2: Multiply H 0 by matrix A s to obtain the weighted embedding H 1 ∈R r×l ;

[0069] Step 4.5.3: Calculate the average value and the maximum value of the second dimension of H 1 , and splice the average value and the maximum value of the second dimension of H 1 as a low-dimensional embedding to H 1 ;

[0070] Step 4.5.4: Splice the low-dimensional embedding with the hidden states at the enhancer and promoter positions to obtain the aggregated feature H ∈ R 4h ;

[0071] The formula for matrix A s is:

[0072]

[0073] In formula (12), A s ∈R r×l , for matrix A s the sum of the elements in each row is 1, and each element in each row is a set of attention coefficients at each position; 0

[0074] The aggregated feature H ∈ R 4h is calculated as follows:

[0075] H = AvgPool(H 1 ) || MaxPool(H 1 ) || h e || h p (13).

[0076] The beneficial effects of the present invention are as follows:

[0077] 1. The present invention proposes the KansformerEPI model. Through an encoder that fuses KAN and Transformer, sequence features and epigenetic features are used to predict enhancer-promoter interactions. When using KAN to integrate features, the non-linear relationship between features is fully considered, the important connections between features are strengthened, and the accuracy of the model in predicting enhancer-promoter interactions is greatly improved.

[0078] 2. The KansformerEPI model of the present invention can achieve excellent prediction performance on multiple mantle cell lines. By comparing with other latest prediction models, the KansformerEPI model shows better performance in multiple evaluation indicators. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 is a flowchart of the method for predicting enhancer-promoter interactions based on multi-omics data provided by the present invention;

[0080] Figure 2 is a schematic structural diagram of the deep neural network model provided by the present invention;

[0081] Figure 3 is a performance comparison diagram of KansformerEPI and other methods for predicting EPIs on the HMEC, IMR90, K562, and NHEK data sets provided by the present invention; Figure 3 In, a is a schematic diagram of the difference in the AUC index between the KansformerEPI model constructed by the present invention and the TransEPI model on four cell lines, and b is a schematic diagram of the difference in the AUPR index between the KansformerEPI model constructed by the invention and the TransEPI model on 4 cell lines;

[0082] Figure 4 This is a comparison graph of the influence of adding a methylation feature set on the prediction results of EPIs provided by the present invention;

[0083] Figure 5 This is a comparison graph of the HMEC, IMR90, K562, and NHEK data sets provided by the present invention before and after balancing the data;

[0084] Figure 6 This is a performance comparison graph of the KansformerEPI model and the TransEPI model after balancing the data provided by the present invention;

[0085] Figure 7 This is the feature importance analysis and window importance analysis of the KansformerEPI model on the GM12878 data set provided by the present invention, Figure 7 where a is a schematic diagram of the relative importance of seven features of the KansformerEPI model on the GM12878 cell line, and b is a schematic diagram of the difference in importance scores of seven features in different windows;

[0086] Figure 8 This is a box plot of the predicted probability distribution of the connection mode between enhancers and promoters on four cell lines provided by the present invention. Detailed implementation manners

[0087] Detailed implementation manner 1: Combining Figure 1 and Figure 2 to illustrate this implementation manner. As Figure 1 shown, the steps of the method for predicting enhancer-promoter interactions based on multi-omics data described in this implementation manner include:

[0088] S1: Obtain multi-omics data and construct a multi-omics data set;

[0089] The multi-omics data includes CTCF binding sites, DNase-I signals, five histone modification features, and DNA methylation information. The multi-omics data set includes Hi-C data sets, CHIA-PET data sets, and feature signals.

[0090] S2: Preprocess the multi-omics data set;

[0091] S201: For the Hi-C dataset and the CHIA-PET dataset, convert enhancer-gene pairs into enhancer-promoter pairs. According to the negative sample generation strategy, pair the measured enhancers with genes that have not been experimentally verified to interact with them. The gene selection criterion is that the distance between the gene and the corresponding enhancer is within 95% of the enhancer-gene distance in the positive samples, and complete the preprocessing of the Hi-C dataset and the CHIA-PET dataset;

[0092] S202: For the feature signals, store the feature signals in bigwig files and narrowPeak files, convert the bigwig files and narrowPeak files into pt files storing tensors, and normalize the features to complete the preprocessing of the feature signals.

[0093] S3: Build a deep neural network model;

[0094] As Figure 2 shown, the deep neural network model adopted in this embodiment is the KansformerEPI model, which consists of a CNN network, a BiLSTM network, and a Kansformer encoder;

[0095] The CNN network is a convolutional neural network, which is used to identify local patterns and key signal features in the input multi-omics dataset;

[0096] The BiLSTM network is a bidirectional long short-term memory network, which is used to perform feature processing from the forward and reverse directions of the multi-omics dataset to enhance the ability to capture time-dependent relationships;

[0097] The Kansformer encoder is used to capture long-range dependencies in the multi-omics dataset. The Kansformer encoder consists of a KAN layer and a Transformer fusion, including an alternating stack of multi-head self-attention layers and the KAN layer. The multi-head self-attention layer is used to capture the relationships between features, and the KAN layer is used for enhancer-promoter interaction relationship prediction.

[0098] S4: Use the preprocessed multi-omics dataset to train the deep neural network model;

[0099] S401: Extract features from the input signal through a convolutional neural network (CNN). The convolutional operation of the CNN layer can effectively identify local patterns and key signal features in the input sequence, providing high-quality feature representations for subsequent layers; The BiLSTM enhances the ability to capture time-dependent relationships through its forward and backward two-layer LSTM networks and outputs the matrix Z.

[0100] S402: The output Z of the BiLSTM is input into the Kansformer encoder, which uses a structure similar to Transformer to capture long-range dependencies. The multi-head self-attention mechanism in the encoder

[31] can capture the relationships between features and calculate query, key, and value vectors, and their calculation formulas are as follows:

[0101] Q = ZW q , K = ZW k , V = ZW v (1);

[0102] In formula (1), the matrix Z ∈ R l×h represents a matrix containing the features of enhancer and promoter sites, and W q , W k and W v respectively represent three weight matrices for mapping changes. K, Q, and V represent three feature encodings after the linear transformation of Z;

[0103] Subsequently, the global relationship matrix M is calculated using K and Q:

[0104]

[0105] In formula (2), d represents the feature dimension of Z. Then, with the help of M, V is feature-enhanced to obtain the enhanced feature

[0106]

[0107] S403: Through the KAN layer, when integrating different features, the non-linear relationship between features is considered. A K-layer KAN network can be represented as the nesting of multiple KAN units:

[0108] KAN(Z) = (Φ K-1 , Φ K-2 , …, Φ 1 , Φ 0 )(Z)(4);

[0109] In formula (4), Φ i is the i-th layer of the KAN network. For each KAN layer, the input dimension is n in , and the output dimension is n out . Each layer of the KAN network is defined as:

[0110]

[0111] For each KAN layer, through the activation function for the input Perform a transformation to obtain the post-activation value Sum up all the post-activation values of the inputs to obtain the activation value of the neuron and the output H of the Kansformer encoder 0 ∈R l×h ;

[0112] Post-activation value The calculation formula is as follows:

[0113]

[0114] S404: The output H from the Kansformer encoder 0 ∈R l×h Use the self-attention embedding method to perform low-dimensional representation learning on H 0 as follows:

[0115] Input H 0 into a two-layer fully connected network:

[0116]

[0117] In formula (7), A s ∈R r×l , for matrix A s , the sum of the elements in each row is 1, and each row element is a set of attention coefficients at each position of H 0 ;

[0118] Multiply H 0 by A s to obtain the weighted embedding H 1 ∈R r×l , then calculate the average value and the maximum value of the second dimension of H 1 , splice them as a low-dimensional embedding to H 1 , finally, splice this low-dimensional embedding with the hidden states at the enhancer and promoter positions to obtain the final aggregated feature H ∈ R 4h ;

[0119] The aggregated feature H ∈ R 4h The calculation formula is as follows:

[0120] H = AvgPool(H 1 ) || MaxPool(H 1 ) || h e || h p (8);

[0121] S405: Input the aggregated H into the KAN classifier to predict the enhancer-promoter interaction probability p. Meanwhile, to make the model sensitive to the positions of enhancers and promoters, the KAN module is also used to predict the distance d between the enhancer and the promoter. p :

[0122] The calculation formula for the enhancer-promoter interaction probability p is:

[0123] p = σ(KAN 1 (H)) * (9);

[0124] In formula (1), σ is the sigmoid function, p is the probability of enhancer-promoter interaction, and the value range of p is 0 to 1.

[0125] The distance d between the enhancer and the promoter p is calculated as:

[0126] d p = KAN 2 (H)(10);

[0127] S406: Based on the calculated enhancer-promoter interaction probability p and the distance d between the enhancer and the promoter p construct a binary cross-entropy loss function for predicting EPIs and a mean squared error loss function for predicting the enhancer-promoter distance Based on the binary cross-entropy loss function and the mean squared error loss function construct a total loss function

[0128] The binary cross-entropy loss function is expressed as:

[0129]

[0130] In formula (11), y i is the true label (0 or 1) of the i-th sample, and p i is the predicted probability;

[0131] The mean squared error loss function is expressed as:

[0132]

[0133] In formula (12), d p is the predicted enhancer-promoter distance, and d t is the true distance;

[0134] The total loss function The expression is:

[0135]

[0136] In formula (13), is the total loss function.

[0137] S407: According to the total loss function train the deep neural network model to obtain the trained deep neural network model;

[0138] S5: Based on the trained deep neural network model, predict the enhancer-promoter interaction relationship.

[0139] In summary, in this embodiment, an encoder integrating KAN and Transformer is used to predict enhancer-promoter interactions using sequence features and epigenetic features. When using KAN to integrate features, the non-linear relationship between features is fully considered, the important connections between features are strengthened, and the accuracy of the model in predicting enhancer-promoter interactions is greatly improved.

[0140] Specific Embodiment 2: In combination with Figures 3 - 8 This embodiment is described. To comprehensively evaluate the effectiveness of the prediction method of this embodiment, the constructed KansformerEPI model is compared with five state-of-the-art deep learning models (TransEPI, TargetFinder, SPEID, 3DPredictor, and DeepTACT models). All models are trained using the same chromosome segmentation scheme to eliminate errors that may be introduced due to differences in data segmentation. The GM12878 and HeLa-S3 datasets are used for training, and the HMEC, IMR90, K562, and NHEK datasets are used for testing.

[0141] This embodiment runs 30 experiments on each cell line and performs exactly the same training and testing operations on its model according to the model parameters provided by TransEPI. To further evaluate the prediction effect of the KansformerEPI model, it is compared with the existing TransEPI model, and the mean difference between the two methods in 4 cell lines is statistically analyzed through a two-sample T-test. This embodiment sets the significance level to 0.05. Under this condition, if the p-value is less than 0.05, it is considered that there is a significant difference in the prediction effects of the two methods on the corresponding cell line; otherwise, if the p-value is greater than 0.05, it is considered that the difference is not significant, as Figure 3As shown in Figure a, the differences in the AUC metrics of the two methods on four cell lines are presented. It can be seen from the results that the KansformerEPI method of this embodiment outperforms TransEPI in all four cell lines, and the differences are statistically significant. Especially in the IMR90 and NHEK cell lines, the AUC of KansformerEPI is increased by approximately 7% and 8%, respectively, while in the HMEC and K562 cell lines, the increase is about 4%. Figure 3 As shown in Figure b, the differences in the AUPR metrics of the two methods on four cell lines are presented. The effects are comparable on the K562 cell line, but there are significant differences on the HMEC, IMR90, and NHEK cell lines, and there are obvious improvements in the effects.

[0142] Table 1 shows the performance comparison of the model of this embodiment with other existing models in terms of AUROC and AUPR metrics. In terms of the AUROC metric, in the HMEC, IMR90, and K562 cell lines, the model of this embodiment is improved by 0.54%, 5.7%, and 3.3%, respectively, compared with the best-performing model among other models. However, in the NHEK cell line, the model of this embodiment is 2.7% lower than the best result among other models. In terms of the AUPR metric, this embodiment finds that in the K562 cell line, the model of this embodiment performs equivalently to the best model among others; while in the HMEC, IMR90, and NHEK cell lines, the model of this embodiment is better than the runner-up model, with improvements of 11.7%, 11.1%, and 4.2%, respectively. Generally speaking, although there is a slight gap in the NHEK cell line, the model of this embodiment performs better than other existing models in most cell lines.

[0143] Table 1

[0144]

[0145] In the experiment of this embodiment, based on the information of CTCF binding sites, DNase-I signals, and five histone modifications, this embodiment further introduces DNA methylation data. After extracting the methylation signals, this embodiment normalizes all features. This embodiment refers to the research of Whalen et al. and obtains the methylation feature sets of four cell lines, namely GM12878, HeLa-S3, IMR90, and K562. These 4 cell lines are selected as the data sets for studying the influence of methylation on the prediction results of EPIs. For training and testing, this embodiment uses the previous scheme to fuse the data of the two cell lines, GM12878 and HeLa-S3, to construct a training set, and then conducts independent tests on the IMR90 and K562 cell lines. During the testing process, this embodiment conducts 30 rounds of experiments respectively, and the results are as Figure 4As shown, in the IMR90 dataset, it can be seen that they are comparable in terms of the AUROC metric, but approximately 1.67% higher in terms of the AUPR metric. In another dataset, K562, both AUROC and AUPR are approximately 1% lower. Although methylation shows a certain promoting effect on EPI prediction in some datasets, opposite results are obtained in other datasets. This difference in results indicates that the effect of methylation on EPI prediction may not be generally applicable. Currently, only these two sets of methylation data are used for testing in this embodiment, so there is a lack of sufficient evidence to prove that methylation has a significant promoting effect on EPI prediction. Although some studies have shown that methylation promotes the prediction of EPIs, this phenomenon is not observed in this embodiment.

[0146] In the task of enhancer-promoter interaction prediction, the positive and negative samples in the original dataset are extremely imbalanced, with some ratios reaching 1:20. The positive samples are determined through wet experiments, while the negative samples are generated by pairing the measured enhancers with genes that do not interact. To observe the impact of data balance on EPI prediction, this paper uses an undersampling strategy in the experimental design to adjust the positive and negative samples of the dataset to 1:1. This embodiment believes that using a 1:1 balanced dataset can effectively reduce the class bias of the model. Since the model in the imbalanced dataset may be biased towards the majority class, i.e., the negative samples, the balanced dataset can make the model pay more attention to the positive samples, thereby improving the recognition ability of the positive samples. This data balance strategy helps to improve the stability of the model's performance on positive samples. As Figure 5 shown, this embodiment conducts a two-sample t-test analysis on the balanced dataset and the imbalanced dataset. It can be seen that the AUPR values of the 4 cell lines on the balanced dataset are significantly higher than those on the imbalanced dataset. Through this adjustment of the dataset, this embodiment can observe that the balanced dataset has a significant improvement in the AUPR metric, and it also helps to evaluate the generalization ability of the model in detecting minority class samples. It can better observe the performance differences of the model on balanced and imbalanced data, and verify whether the model can accurately identify enhancer-promoter interactions when the distribution of positive and negative samples changes. At the same time, this embodiment also compares the performance with other state-of-the-art models on the balanced dataset. The comparison results are as Figure 6 shown, the models of this embodiment are far higher than TransEPI, which shows that the models of this embodiment are also superior to the TransEPI model on the balanced dataset and have higher stability.

[0147] In this embodiment, a random forest model is used to deeply analyze the key features of enhancer-promoter interactions, focusing on the importance of the features and their contributions within different windows. To this end, in this embodiment, the importance scores of each feature in multiple windows are first calculated to evaluate its impact on predicting EPIs. To ensure the systematicness of the analysis, in this embodiment, some relatively important windows of enhancers and promoters are screened out from 5000 windows, and the contributions of 7 features in each window are calculated respectively; through this method of window-by-window analysis, in this embodiment, the changes in the importance of each feature within different gene regulatory regions can be observed, so as to identify which features have a significant impact on predicting EPIs within certain specific windows, such as Figure 7 as shown in a, in this embodiment, the importance scores of the features within these windows are calculated and their average values are taken to obtain the relative importance of the seven features in the GM12878 cell line. The analysis results show that the DNase-I signal has the highest score, indicating that its importance in prediction is the most prominent. In contrast, the CTCF binding site has the lowest score, while the two features of H3K4me1 and H3K4me3 have relatively high importance. The possible reason is that these two features usually appear in enhancer regions, and enhancers are relatively active in the prediction of EPIs.

[0148] Subsequently, in this embodiment, a heatmap is used to visualize the distribution of the importance scores of the features within each window, as Figure 7 shown in b. The heatmap shows the differences in the importance scores of 7 features (including CTCF, DNase-I, and five histone modification signals) within different windows. The depth of the color in the heatmap reflects the high or low importance score of the feature. The darker the color, the higher the importance of the feature in predicting EPIs, and the lighter the color, the lower the importance. Through heatmap analysis, in this embodiment, the contribution of each feature within different genomic regions can be understood more intuitively. e and p represent enhancer and promoter windows, and windows such as e-1 and p-1 represent the windows near them. Within the enhancer window, in this embodiment, it can be seen that the color of H3K4me1 is the darkest, and this feature often appears in enhancer regions. In the windows near the enhancer, it can be found that the features of H3K9me3, H3K27me3, and CTCF have higher importance. Similarly, it can be found that the feature of H3K9me3 has the darkest color and is more important within the promoter window; within the windows near the promoter, the two features of H3K4me1 and H3K4me3 have higher importance. This observation provides new biological insights into the regulatory mechanism of EPI, further verifying the key role of these features in gene regulation; in addition, these results also provide an important reference basis for subsequent more in-depth biological research, helping this embodiment better understand EPI and its biological functions.

[0149] Table 2

[0150]

[0151] Enhancer-promoter interactions usually form subnetworks of various biological processes. To better analyze and understand these interaction patterns, in this embodiment, the number of promoters regulated by each enhancer (out-degree) and the number of enhancers regulating each promoter (in-degree) in the positive sample were calculated. In the data of this embodiment, on average, each enhancer regulates approximately 1 to 2 promoters, while a single promoter is regulated by 2 - 3 enhancers. This embodiment was inspired by Sushmita Roy

[32] et al., and classified the connection patterns between these enhancers and promoters into the following four categories: (i) Single interaction (Single): that is, a one-to-one relationship, where one enhancer interacts with only one promoter; (ii) Multi-output component (MO): that is, one enhancer regulates multiple promoters; (iii) Multi-input component (MI): that is, one promoter is simultaneously regulated by multiple enhancers; (iv) Multi-input multi-output component (MIMO): that is, multiple enhancers regulate multiple promoters, forming a network; When analyzing 4 test cell lines, this embodiment counted the number of these four regulatory relationships and the predicted probabilities in the model of this embodiment. The results showed that among these 4 regulatory relationships, the proportion of MI was the largest, as shown in Table 2. At the same time, through the Figure 8 box plot, this embodiment observed the distribution of predicted values of the 4 regulatory relationships and found that the predicted probabilities of MI were basically optimal. This result indicates that in the dataset BENGI, this association relationship of MI not only has an advantage in quantity but also is the best in terms of prediction accuracy; Combining the above examples, it confirms the excellent prediction performance of KansformerEPI.

[0152] The above is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Although the present invention has been disclosed above with a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the relevant art can make some changes or modifications within the scope of the technical solution of the present invention to make equivalent embodiments with equivalent changes. However, as long as it does not depart from the content of the technical solution of the present invention and is based on the technical essence of the present invention, any simple modification, equivalent replacement, and improvement of the above embodiments within the spirit and principle of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A method for predicting enhancer-promoter interactions based on multi-omics data, characterized in that: The method for predicting enhancer-promoter interactions based on multi-omics data comprises the following steps: Step 1: Build a deep neural network model; Step 2: Acquire multi-omics data and construct multi-omics datasets; Step 3: preprocessing the multi-omics dataset; Step 4: Train the deep neural network model based on the preprocessed multi-omics dataset; Step 5: Predict enhancer-promoter interactions based on the trained deep neural network model.

2. The method for predicting enhancer-promoter interactions based on multi-omics data according to claim 1, characterized in that: The deep neural network model in step 1 consists of a CNN network, a BiLSTM network, and a Kansformer encoder; The CNN network is a convolutional neural network used to identify local patterns and key signal features in input multi-omics data sets; The BiLSTM network is a bidirectional long short-term memory network, which is used to process features from the forward and reverse directions of multi-omics data sets and enhance the ability to capture time dependencies; The Kansformer encoder is used to capture long-range dependencies in multi-omics datasets.

3. The method for predicting enhancer-promoter interactions based on multi-omics data according to claim 2, characterized in that: The Kansformer encoder consists of a fusion of KAN layers and Transformer, including alternating stacking of multi-head self-attention layers and KAN layers. The multi-head self-attention layers are used to capture the relationship between features, and the KAN layers are used to predict enhancer-promoter interaction relationships.

4. The method for predicting enhancer-promoter interactions based on multi-omics data according to claim 1, characterized in that: The multi-omics data in step 2 include CTCF binding sites, DNase-I signals, five histone modification features and DNA methylation information, and the multi-omics datasets include Hi-C datasets, CHIA-PET datasets and characteristic signals.

5. The method for predicting enhancer-promoter interactions based on multi-omics data according to claim 1, characterized in that: Step 3 specifically includes: Step 3.1: For Hi-C datasets and CHIA-PET datasets, the enhancer-gene pairs were converted into enhancer-promoter pairs. According to the negative sample generation strategy, the measured enhancers were selected to be paired with genes that have not been experimentally verified to interact with them. The gene selection criterion was that the distance between the enhancer and the corresponding enhancer was within 95% of the positive sample enhancer-gene distance, completing the preprocessing of the Hi-C dataset and CHIA-PET dataset; Step 3.2: For the feature signal, store the feature signal in a bigwig file and a narrowPeak file, convert the bigwig file and the narrowPeak file into a pt file storing tensors, normalize the features, and complete the preprocessing of the feature signal.

6. The method for predicting enhancer-promoter interactions based on multi-omics data according to claim 1, characterized in that: Step 4 specifically includes: Step 4.1: The preprocessed multi-omics dataset is used as an input sequence, and features are extracted from the input sequence through the CNN network; Step 4.2: The BiLSTM network receives the extracted features and enhances the ability of the extracted features to capture time dependencies through the forward and backward LSTM networks, outputting the matrix Z, Z∈R l×h is a matrix containing the features of enhancer and promoter sites, where R is a real number field, l is the number of rows in the matrix, corresponding to the number of sequences, and h is the number of columns in the matrix, corresponding to the dimension of the features; Step 4.3: Input the output matrix Z of the BiLSTM network into the Kansformer encoder, and capture the relationship between the features in the output matrix Z through the multi-head self-attention layer to obtain the enhanced features Step 4.4: Enhanced features Input the KAN layer for feature integration and get the output H0∈R of the Kansformer encoder l×h ; Step 4.5: Use the self-attention embedding method to learn the low-dimensional representation of H0 and obtain the aggregated feature H∈R 4h ; Step 4.6: Input the aggregated features H into the KAN classifier to calculate the probability p of predicting enhancer-promoter interaction and the distance d between enhancers and promoters. p ; Step 4.7: Based on the calculated enhancer-promoter interaction probability p and the distance d between the enhancer and the promoter p Constructing a binary cross entropy loss function for EPIs prediction And the mean square error loss function for predicting enhancer-promoter distance Based on binary cross entropy loss function And the mean square error loss function Constructing the total loss function Step 4.8: According to the total loss function Train the deep neural network model to obtain a trained deep neural network model; The calculation formula for enhancer-promoter interaction probability p is: p = σ(KAN1(H))*(1); In formula (1), σ is the sigmoid function, p is the probability of enhancer-promoter interaction, and the p value ranges from 0 to 1; The distance between enhancer and promoter p The calculation formula is: d p =KAN2(H) (2); Binary Cross Entropy Loss Function The expression is: In formula (3), y i is the true label (0 or 1) of the i-th sample, p i is the predicted probability; Mean Squared Error Loss Function The expression is: In formula (4), d p is the predicted enhancer-promoter distance, d t is the real distance; Total loss function The expression is: In formula (5), is the total loss function.

7. The method for predicting enhancer-promoter interactions based on multi-omics data according to claim 6, characterized in that: Step 4.3 specifically includes: Step 4.3.1: Calculate the query Q, key Q and value vector V based on the matrix Z and the weight matrix used to map the changes; Step 4.3.2: Calculate the global relationship matrix M based on the query Q and key Q; Step 4.3.3: Enhance the value vector V through the global relationship matrix M to obtain the enhanced features The calculation formula for query Q, key Q and value vector V is: Q=ZW q ,K=ZW k ,V=ZW v (6); In formula (6), W q , W k and W v are three weight matrices used to map changes, K, Q, and V are three feature codes after the matrix Z is linearly transformed; The calculation formula of the global relationship matrix M is: In formula (7), d is the characteristic dimension of matrix Z; Enhanced features The calculation formula is: In formula (8), b is the deviation.

8. The method for predicting enhancer-promoter interactions based on multi-omics data according to claim 6, characterized in that: The KAN layer in step 4.4 is a K-layer KAN network. A K-layer KAN network includes nested KAN units of multiple layers. The steps of feature integration include: Step 4.4.1: For each KAN layer, pass the activation function For input After transformation, the activation value is obtained Step 4.4.2: Sum the post-activation values ​​of all inputs to obtain the activation value of the neuron and the output H0∈R of the Kansformer encoder l×h ; The expression of KAN network is: KAN(Z)=(Φ K-1 ,F K-2 ,…,Φ1,Φ0)(Z) (9); In formula (9), Φ i is the i-th layer of the KAN network. For each KAN layer, the input dimension is n in , the output dimension is n out , each layer of KAN network is defined as: Post-activation value The calculation formula is:

9. The method for predicting enhancer-promoter interactions based on multi-omics data according to claim 6, characterized in that: Step 4.5 specifically includes: Step 4.5.1: Input H0 into a two-layer fully connected network to obtain the matrix A s ; Step 4.5.2: Compare H0 with the matrix A s Multiply them together to get the weighted embedding H1∈R r×l ; Step 4.5.3: Calculate the average and maximum value of the second dimension of H1, and concatenate the average and maximum value of the second dimension of H1 into H1 as a low-dimensional embedding; Step 4.5.4: Concatenate the low-dimensional embedding with the hidden states of the enhancer and promoter positions to obtain the aggregated feature H∈R 4h : Matrix A s The calculation formula is: In formula (12), A s ∈R r×l , matrix A s The sum of each row of elements is 1, and each row of elements is a set of attention coefficients at each position of H0; The aggregated feature H∈R 4h The calculation formula is: H=AvgPool(H1)||MaxPool(H1)||h e ||h p (13)。

Citation Information

Patent Citations

  • Chromatin topological correlation domain boundary prediction method based on multi-modal fusion

    CN115831217A

  • Multi-head attention mechanism-based enhancer-promoter interaction prediction model construction method

    CN116312748A

  • Construction method and application of double sgRNA library

    CN117947525A

  • Joint modeling method and apparatus for enhancing local features of pedestrians

    WO2024060321A1

Cited By

  • Neural network calculation method and device for gene expression regulation and control analysis

    CN121306251A

  • A neural network computing method for gene expression regulation analysis and device thereof

    CN121306251B