A method for predicting promoter-enhancer interaction based on DNA sequence

By employing a DNA sequence-based promoter-enhancer interaction prediction method, which combines windowing and k-mer processing with adaptive feature fusion and cross-attention computation, this approach addresses the shortcomings of existing models in feature extraction and prediction accuracy. It achieves higher-precision EPI prediction, supporting the identification of disease-related regulatory elements and personalized treatment.

CN120072055BActive Publication Date: 2026-01-27XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510224961.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2026-01-27
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

Existing promoter-enhancer interaction prediction models have significant shortcomings in feature extraction, feature fusion, and prediction accuracy, resulting in low accuracy and reliability of EPI prediction.

Method used

A promoter-enhancer interaction prediction method based on DNA sequences is adopted. Sequence fragments are extracted by truncating them through a fixed-length window, and after k-mer processing, global and multi-scale feature extraction is performed using a promoter-enhancer prediction model. Combined with feature adaptive fusion and cross-attention calculation, an interactive modeling mechanism is established to improve prediction accuracy.

Benefits of technology

It significantly improves the accuracy and reliability of promoter-enhancer interaction prediction, enabling more accurate capture of their interaction patterns, providing potential drug targets and a basis for personalized treatment, and promoting the development of functional genomics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072055B_ABST
    Figure CN120072055B_ABST
Patent Text Reader

Abstract

The present application relates to a DNA sequence-based promoter-enhancer interaction prediction method, comprising: using a window to intercept the obtained multiple groups of promoter sequences and multiple groups of enhancer sequences to obtain multiple groups of promoter sequence fragments and multiple groups of enhancer sequence fragments; each group of promoter sequence fragments and each group of enhancer sequence fragments are subjected to k-mer processing to obtain multiple groups of k-mer processed promoter sequence fragments and enhancer sequence fragments; global feature extraction and multi-scale feature extraction are performed on multiple groups of k-mer processed promoter sequence fragments and enhancer sequence fragments to obtain promoter global representation, first multi-scale feature, enhancer global representation and second multi-scale feature; then the obtained multiple features are subjected to feature adaptive fusion and cross attention calculation to obtain cross attention results; based on the cross attention results, the prediction results are calculated. The method can effectively improve the model prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bioinformatics technology, specifically relating to a method for predicting promoter-enhancer interactions based on DNA sequences. Background Technology

[0002] Enhancers are important non-coding regulatory elements in the genome that can regulate gene expression over long distances by interacting with promoters. The enhancer-promoter interaction (EPI) is a key mechanism in the gene expression regulatory network, playing a crucial role in cell differentiation, organ development, and disease development.

[0003] Enhancers activate the promoters of their target genes through specific transcription factor binding sites, promoting RNA polymerase binding and transcription initiation. This interaction between enhancers and promoters is not limited to neighboring genes within a linear genomic distance but can also affect genes at greater distances through three-dimensional genomic conformation. The function of enhancer-promoter interactions (EPIs) is crucial for cell development, embryonic morphogenesis, and the regulation of the immune system. Studies have shown that abnormalities in enhancer-promoter interactions may lead to a variety of complex diseases, including cancer, autoimmune diseases, and neurodegenerative diseases. Therefore, enhancing-promoter interaction prediction can help identify key regulatory elements associated with diseases and reveal the underlying mechanisms of abnormal gene expression. For example, predicting enhancers associated with cancer-driving genes can elucidate the transcriptional regulatory mechanisms of tumorigenesis. EPI prediction can also provide potential drug targets, particularly those with inactivated or abnormally enhanced enhancers, providing a basis for personalized treatment. Reconstructing genomic regulatory networks can improve our understanding of complex phenotypes and polygenic traits, advancing the field of functional genomics.

[0004] Although the biological importance of enhancer-promoter interactions is widely recognized, accurately predicting these interactions remains a significant challenge. Several existing promoter-enhancer interaction prediction models exist, but these models still have considerable shortcomings in feature extraction, feature fusion, and prediction accuracy, severely reducing the precision and reliability of EPI predictions. Summary of the Invention

[0005] To address the aforementioned problems in the existing technology, this invention provides a method for predicting promoter-enhancer interactions based on DNA sequences. The technical problem to be solved by this invention is achieved through the following technical solution:

[0006] This invention provides a method for predicting promoter-enhancer interactions based on DNA sequences, comprising:

[0007] Using a fixed-length window, multiple sets of promoter sequences and multiple sets of enhancer sequences are truncated to obtain multiple sets of promoter sequence fragments and multiple sets of enhancer sequence fragments. K-mer processing is then applied to each set of promoter sequence fragments and enhancer sequence fragments to obtain multiple sets of K-mer processed promoter sequence fragments and multiple sets of K-mer processed enhancer sequence fragments. Based on the promoter-enhancer prediction model, global feature extraction and multi-scale feature extraction are performed on the multiple sets of K-mer processed promoter sequence fragments to obtain a global promoter representation and a first multi-scale feature. Similarly, global feature extraction and multi-scale feature extraction are performed on the multiple sets of K-mer processed enhancer sequence fragments to obtain a global enhancer representation and a second multi-scale feature. Feature adaptive fusion and cross-attention calculation are then performed sequentially on the global promoter representation, the global enhancer representation, the first multi-scale feature, and the second multi-scale feature to obtain a cross-attention result. Based on the cross-attention result, the prediction result is calculated.

[0008] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0009] To address the significant shortcomings of existing promoter-enhancer interaction prediction models in feature extraction, feature fusion, and prediction accuracy, which severely reduces the precision and reliability of EPI prediction, this invention provides a DNA sequence-based promoter-enhancer interaction prediction method. This method utilizes a promoter-enhancer prediction model to perform global and multi-scale feature extraction on multiple sets of k-mer-processed promoter sequence fragments, capturing information from the DNA sequence at multiple levels and scales. Through adaptive feature fusion, the contribution of features at different scales is adjusted to fully consider the importance of different features and fuse feature information of varying scales and complexities. Subsequently, cross-attention computation is used to establish an effective interaction modeling mechanism, accurately capturing the interaction patterns between promoters and enhancers. Finally, the prediction results are calculated using the cross-attention results, significantly improving the model's predictive ability for promoter-enhancer interactions and effectively enhancing the model's prediction accuracy and reliability. Attached Figure Description

[0010] Figure 1 This is a schematic flowchart of a promoter-enhancer interaction prediction method based on DNA sequence provided in an embodiment of the present invention;

[0011] Figure 2 This is an application example diagram of the promoter-enhancer interaction prediction method based on DNA sequence provided in the embodiments of the present invention. Detailed Implementation

[0012] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0013] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0014] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.

[0015] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings, the disclosure, and the appended claims in carrying out the claimed invention. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.

[0016] The present invention will now be described in detail with reference to the accompanying drawings, a method for predicting promoter-enhancer interactions based on DNA sequences.

[0017] Figure 1 This is a schematic flowchart of a promoter-enhancer interaction prediction method based on DNA sequences provided in an embodiment of the present invention. Figure 1 As shown, the method includes steps 110-150. Specifically:

[0018] Step 110: Using a fixed-length window, truncate the obtained multiple sets of promoter sequences and multiple sets of enhancer sequences to obtain multiple sets of promoter sequence fragments and multiple sets of enhancer sequence fragments.

[0019] Here, step 110 specifically includes: using a window to slide and truncate each of the multiple sets of promoter sequences to obtain multiple sets of promoter sequence fragments; using a window to slide and truncate each of the multiple sets of enhancement sequences to obtain multiple sets of enhancement sequence fragments.

[0020] For example, multiple sets of promoter sequences P and multiple sets of enhancer sequences E belong to publicly available validated promoter-enhancer interaction datasets from the human genome. Furthermore, one set of promoter sequences P and one set of enhancer sequences E constitute a promoter-enhancer sequence pair. The number of promoter sequences P and enhancer sequences E are the same, as are the number of fragments obtained by truncating both. Here, the window length can be 2 kbp or 3 kbp. It should be understood that the present invention does not limit the size of the window and can adjust it according to actual conditions.

[0021] Step 120: Perform k-mer processing on each group of promoter sequence fragments and each group of enhancer sequence fragments to obtain multiple groups of k-mer processed promoter sequence fragments and multiple groups of k-mer processed enhancer sequence fragments.

[0022] Here, the core of k-mer processing is to use the sliding window technique to map each small segment (of length k) in the DNA sequence to a numerical representation. Taking k=6 as an example, for the sequence ACGTAGCTG, k-mer processing will convert it into four fragments [ACGTAG, CGTAGC, GTAGCT, TAGCTG], and each fragment will be converted into a vector of fixed length.

[0023] Here, k-mer processing is used to further divide the fixed-length sequence segment, enabling the promoter sequence P and enhancer sequence E to be used as input data in subsequent deep learning models. Typically, k takes a value between 5 and 10; shorter k values ​​help capture local sequence features, while longer k values ​​capture more global information. For example, suppose the fixed-length promoter sequence P is represented as... Therefore, the promoter sequence fragment after k-mer processing can be represented as: Similarly, a fixed-length enhancer subsequence E is represented as... Therefore, the promoter sequence fragment after k-mer processing can be represented as:

[0024] Step 130: Based on the promoter-enhancer prediction model, perform global feature extraction and multi-scale feature extraction on multiple groups of k-mer processed promoter sequence fragments to obtain global promoter representation and first multi-scale feature, and perform global feature extraction and multi-scale feature extraction on multiple groups of k-mer processed enhancer sequence fragments to obtain global enhancer representation and second multi-scale feature.

[0025] Specifically, the promoter-enhancer prediction model includes a first granularity feature extraction module and a second granularity feature extraction module; step 130 includes: (1) using the first granularity feature extraction module to perform global feature extraction on multiple sets of k-mer processed promoter sequence fragments to obtain global promoter representation, and to perform global feature extraction on multiple sets of k-mer processed enhancer sequence fragments to obtain global enhancer representation; (2) using the second granularity feature extraction module to perform multi-scale feature extraction on multiple sets of k-mer processed promoter sequence fragments to obtain first multi-scale features, and to perform multi-scale feature extraction on multiple sets of k-mer processed enhancer sequence fragments to obtain second multi-scale features.

[0026] It should be noted that the granularity of the first granularity feature extraction module is greater than that of the second granularity feature extraction module. In other words, the first granularity feature extraction module can be understood as a coarse-grained feature extraction module, and the second granularity feature extraction module can be understood as a fine-grained feature extraction module.

[0027] Here, step (1) further includes: inputting each group of k-mer processed promoter sequence fragments and each group of k-mer processed enhancer sequence fragments into the pre-trained Nucleotide Transformers model to obtain the corresponding first context-aware representation sequence and second context-aware representation sequence; performing global max pooling and global average pooling on each group of first context-aware representation sequences to obtain the corresponding first max pooling data and first average pooling data, and performing global max pooling and global average pooling on each group of second context-aware representation sequences to obtain the corresponding second max pooling data and second average pooling data; concatenating the first max pooling data and first average pooling data corresponding to multiple groups of k-mer processed promoter sequence fragments to obtain the global representation of the promoter; and concatenating the second max pooling data and second average pooling data corresponding to multiple groups of k-mer processed enhancer sequence fragments to obtain the global representation of the enhancer.

[0028] Here, context-aware sequence representation refers to the ability of a model to understand and utilize the contextual information surrounding each element in a sequence during sequence modeling, thereby more accurately predicting or generating sequence content. The Nucleotide Transformers model refers to a transformer model based on amino acid gene sequences. The pre-trained Nucleotide Transformers model uses the Rotated Position Encoding (RoPE) algorithm to process each set of k-mer processed promoter sequence fragments and each set of k-mer processed enhancer sequence fragments. The RoPE algorithm leverages the advantages of sinusoidal frequency distribution in position encoding, combining it with traditional sine wave methods. Specifically, for each position, RoPE calculates its frequency range and then uses the frequency values ​​for complex rotation. The frequency calculation is position-dependent, similar to traditional sine wave encoding, but in RoPE, the position vector is rotated in each dimension, thus capturing relative positional information. During the rotation, each token rotates by a different angle in different dimensions, representing the relative position of that location. RoPE effectively combines the rotational properties of sine and cosine functions, enabling it to better represent relative positional information. Specifically, RoPE computes a set of rotation matrices for each position and then uses them to transform the input token embedding. This transformation helps the model learn the relationships between adjacent tokens and their global relative positions.

[0029] Furthermore, compared to traditional positional encoding, which uses fixed sine and cosine wave functions to represent the uniqueness of each position and assigns an absolute positional identifier to each position—an encoding that is fixed and does not change with training data—RoPE does not directly provide the absolute value of each position. Instead, it captures the relative positional relationship through rotational encoding. This approach gives RoPE greater flexibility and scalability, maintaining good performance on inputs of varying lengths.

[0030] In one possible implementation, the first context-aware representation sequence H P It can be represented as: H P =Nucleotide Transformers(X P ), and the second context-aware representation sequence H E It can be represented as: H E =Nucleotide Transformers(X E ).

[0031] It's important to note that global max pooling and global average pooling are not performed in any particular order. Global max pooling extracts the maximum value for each feature dimension in the sequence, while global average pooling averages the values ​​across all feature dimensions. The information extracted by max pooling and average pooling is complementary. Max pooling focuses on salient features, while average pooling focuses on the overall features of the sequence. Concatenating the results of both preserves both the strongest signal and global information, resulting in a richer and more comprehensive feature representation.

[0032] In one possible implementation, the global promoter representation P global It can be represented as: P global =Concat(MaxPool(H P ),AvgPool(H P )); where MaxPool(·) is global max pooling, AvgPool(·) is global average pooling, and GlobalFeatureConcat is global feature concatenation.

[0033] In one possible implementation, the global representation of the enhancer can be expressed as: E global =Concat(MaxPool(H E ),AvgPool(H E )).

[0034] Here, step (2) specifically includes: using multiple convolutional networks of different scales, extracting features from each group of k-mer processed promoter sequence fragments and each group of k-mer processed enhancer sequence fragments to obtain multiple promoter convolutional features and multiple enhancer convolutional features; performing batch normalization and ReLU activation on the multiple promoter convolutional features and multiple enhancer convolutional features to obtain optimized promoter convolutional features and optimized enhancer convolutional features; using a convolutional kernel of size 1, performing secondary convolution on the optimized promoter convolutional features and optimized enhancer convolutional features to obtain the first multi-scale feature and the second multi-scale feature.

[0035] For example, multiple convolutional networks with different scales include convolutional kernels of size 3, size 5, and size 7 processed in parallel. Multi-scale local features are extracted by using convolutional layers with different kernel sizes. Here, convolutional kernels of size 3, 5, and 7 are used to extract features from each group of k-mer processed promoter sequence fragments and each group of k-mer processed enhancement sequence fragments, respectively. After feature extraction, batch normalization and ReLU activation are used to enhance nonlinear expressive power and reduce internal covariance shift.

[0036] In one possible implementation, the computation process of each convolutional kernel can be represented as: ConV(x) k =ReLU(W k *x+b k ), where W k It is a convolution kernel, b k is the bias term, whose parameters change automatically as the model trains; * is the convolution operation.

[0037] Step 140: Perform feature adaptive fusion and cross-attention calculation on the global representation of the promoter, the global representation of the enhancer, the first multi-scale feature, and the second multi-scale feature in sequence to obtain the cross-attention result.

[0038] Here, step 140 specifically includes: (a) using a gating mechanism to perform feature adaptive fusion on the global representation of the promoter, the global representation of the enhancer, the first multi-scale feature, and the second multi-scale feature to obtain the promoter fusion feature and the enhancer fusion feature; (b) using the promoter fusion feature and the enhancer fusion feature to construct attention key-value pairs to represent the similarity between the global representation of the promoter and the global representation of the enhancer; and (c) using the attention key-value pairs to calculate the cross-attention result.

[0039] Step (a) specifically includes: obtaining the first gating weight, the second gating weight, and the third gating weight; performing calculations on the first gating weight and the global representation of the promoter to obtain the first gating data; performing calculations on the second gating weight and the first multi-scale feature to obtain the second gating data; and using the sigmoid activation function to perform calculations on the third gating weight, the global representation of the promoter, and the first multi-scale feature to obtain the third gating data; performing weighted calculations on the first gating data, the second gating data, and the third gating data to obtain the promoter fusion feature; performing calculations on the first gating weight and the global representation of the enhancer to obtain the fourth gating data; performing calculations on the second gating weight and the second multi-scale feature to obtain the fifth gating data; and using the sigmoid activation function to perform calculations on the third gating weight, the global representation of the enhancer, and the second multi-scale feature to obtain the sixth gating data; and performing weighted calculations on the fourth, fifth, and sixth gating data to obtain the enhancer fusion feature. It should be noted that the first gating weight, the second gating weight, and the third gating weight are obtained using a deep learning network.

[0040] In one possible implementation, the promoter fusion feature P fused It can be represented as:

[0041] P fused =z·h g +(1-z)·h l ;

[0042] h g =tanh(W g ·P global );

[0043] h l =tanh(W l ·P local );

[0044] z=σ(W z ·[P global ,P local ]);

[0045] Among them, P local This refers to the first multi-scale feature, W. g It is the first gating weight, W l It is the second gating weight, W z σ is the third gating weight, and σ is the sigmoid activation function.

[0046] In one possible implementation, the enhanced sub-fusion feature E fused It can be represented as:

[0047] E fused =z·h g +(1-z)·h l ;

[0048] h g =tanh(W g ·E global );

[0049] h l =tanh(W l ·E local );

[0050] z=σ(W z ·[E global E local ]);

[0051] Among them, E local This refers to the second multi-scale feature.

[0052] Here, the attention key-value pairs in step (b) include: a query matrix Q responsible for mapping facilitator features to the query space. P Key matrix K E Value matrix V E , where the query matrix Q P The expression is: Q P =P fused W Q Key matrix K E The expression is: K E =Efused W K Value matrix V E The expression is: V E =E fused W V .

[0053] Here, the cross-attention result in step (c) is expressed as: CrossAtten(P,E)=A·V E Where A is the attention weight matrix, representing the similarity between promoter and enhancer features, and its expression is: Here, D refers to the attention scaling parameter.

[0054] Step 150: Calculate the prediction result based on the cross-attention result.

[0055] Here, step 150 specifically includes: inputting the cross-attention result into a multilayer perceptron to calculate the prediction result. The multilayer perceptron comprises several fully connected layers. A multilayer perceptron (MLP) is a feedforward artificial neural network composed of multiple neurons (neural nodes) arranged hierarchically, including an input layer, hidden layers, and an output layer. Neurons between layers are connected by weights, and information propagates sequentially from the input layer to the output layer without feedback connections. After computation using each layer in the MLP, to avoid overfitting, the output value of each layer is processed using layer normalization (LayerNorm) and dropout.

[0056] Here, the multilayer perceptron uses a binary cross-entropy loss function to constrain the output value, and its formula is: Loss=-∑(y true *log(y pred )+(1-y true )*log(1-y pred Furthermore, the formula for calculating the output value Y of the multilayer perceptron is Y = MLP(CrossAtten(P,E)), and the prediction result can be expressed as: y pred =σ(Y).

[0057] Figure 2 This is an application example diagram of the promoter-enhancer interaction prediction method based on DNA sequences provided in this embodiment of the invention. For example... Figure 2As shown, multiple promoter-enhancer sequences are processed by standardization (i.e., truncating using a fixed-length window) and k-mer processing, and then input into the first and second granularity feature extraction modules in the promoter-enhancer prediction model, respectively. The first granularity feature extraction module extracts the overall feature information of the sequence, obtaining the global representation of the promoter-enhancer. The second granularity feature extraction module extracts the local feature information of the sequence from multiple levels and scales, obtaining multiple sets of local features. Subsequently, a gating unit adaptively fuses the multiple input feature information based on actual needs and the importance of each feature, and a cross-attention unit performs interactive calculations on the fused feature information. Finally, the prediction unit predicts the interactive calculation results to obtain the final prediction result.

[0058] To address the significant shortcomings of existing promoter-enhancer interaction prediction models in feature extraction, feature fusion, and prediction accuracy, which severely reduces the precision and reliability of EPI prediction, this invention provides a DNA sequence-based promoter-enhancer interaction prediction method. This method utilizes a promoter-enhancer prediction model to perform global and multi-scale feature extraction on multiple sets of k-mer-processed promoter sequence fragments, capturing information from the DNA sequence at multiple levels and scales. Through adaptive feature fusion, the contribution of features at different scales is adjusted to fully consider the importance of different features and fuse feature information of varying scales and complexities. Subsequently, cross-attention computation is used to establish an effective interaction modeling mechanism, accurately capturing the interaction patterns between promoters and enhancers. Finally, the prediction results are calculated using the cross-attention results, significantly improving the model's predictive ability for promoter-enhancer interactions and effectively enhancing the model's prediction accuracy and reliability.

[0059] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. A method for predicting promoter-enhancer interactions based on DNA sequences, characterized in that, include: Using a fixed-length window, the acquired multiple sets of promoter sequences and multiple sets of enhancement sequences are truncated to obtain multiple sets of promoter sequence fragments and multiple sets of enhancement sequence fragments; Each group of promoter sequence fragments and each group of enhancer sequence fragments were processed by k-mer to obtain multiple groups of k-mer processed promoter sequence fragments and multiple groups of k-mer processed enhancer sequence fragments. Based on the promoter-enhancer prediction model, global feature extraction and multi-scale feature extraction are performed on the multiple sets of k-mer processed promoter sequence fragments to obtain global promoter representation and first multi-scale feature; and global feature extraction and multi-scale feature extraction are performed on the multiple sets of k-mer processed enhancer sequence fragments to obtain global enhancer representation and second multi-scale feature. The global representation of the promoter, the global representation of the enhancer, the first multi-scale feature, and the second multi-scale feature are sequentially subjected to feature adaptive fusion and cross-attention calculation to obtain the cross-attention result. Based on the cross-attention results, the prediction results are calculated; The step of sequentially performing feature adaptive fusion and cross-attention calculation on the global representation of the promoter, the global representation of the enhancer, the first multi-scale feature, and the second multi-scale feature to obtain the cross-attention result includes: Using a gating mechanism, the global representation of the promoter, the global representation of the enhancer, the first multi-scale feature, and the second multi-scale feature are adaptively fused to obtain the promoter fusion feature and the enhancer fusion feature. Using the promoter fusion features and enhancer fusion features, attention key-value pairs are constructed to characterize the similarity between the global representation of the promoter and the global representation of the enhancer; The cross-attention result is calculated using the attention key-value pairs; The prediction result calculated based on the cross-attention result includes: The cross-attention result is input into a multilayer perceptron to calculate the prediction result, wherein the multilayer perceptron includes several fully connected layers.

2. The method for predicting promoter-enhancer interactions based on DNA sequences according to claim 1, characterized in that, The promoter-enhancer prediction model includes a first granularity feature extraction module and a second granularity feature extraction module; The promoter-enhancer prediction model performs global feature extraction and multi-scale feature extraction on the multiple sets of k-mer processed promoter sequence fragments to obtain a global promoter representation and a first multi-scale feature. It also performs global feature extraction and multi-scale feature extraction on the multiple sets of k-mer processed enhancer sequence fragments to obtain a global enhancer representation and a second multi-scale feature, including: Using the first granularity feature extraction module, global feature extraction is performed on the multiple sets of k-mer processed promoter sequence fragments to obtain the global representation of the promoter, and global feature extraction is performed on the multiple sets of k-mer processed enhancer sequence fragments to obtain the global representation of the enhancer. Using the second granularity feature extraction module, multi-scale feature extraction is performed on the multiple sets of k-mer processed promoter sequence fragments to obtain the first multi-scale feature, and multi-scale feature extraction is performed on the multiple sets of k-mer processed enhancer sequence fragments to obtain the second multi-scale feature.

3. The method for predicting promoter-enhancer interactions based on DNA sequences according to claim 2, characterized in that, The step of using the first granularity feature extraction module to perform global feature extraction on the multiple sets of k-mer processed promoter sequence fragments to obtain the global representation of the promoter, and to perform global feature extraction on the multiple sets of k-mer processed enhancer sequence fragments to obtain the global representation of the enhancer, includes: Each set of k-mer processed promoter sequence fragments and each set of k-mer processed enhancer sequence fragments are input into the pre-trained Nucleotide Transformers model to obtain the corresponding first context-aware representation sequence and second context-aware representation sequence. Global max pooling and global average pooling are performed on each group of first context-aware representation sequences to obtain the corresponding first max pooling data and first average pooling data. Similarly, global max pooling and global average pooling are performed on each group of second context-aware representation sequences to obtain the corresponding second max pooling data and second average pooling data. The first max pooling data and the first average pooling data corresponding to the multiple sets of k-mer processed promoter sequence fragments are concatenated to obtain the global representation of the promoter. The second max pooling data and the second average pooling data corresponding to the multiple sets of k-mer processed enhancer sequence fragments are concatenated to obtain the global representation of the enhancer.

4. The method for predicting promoter-enhancer interactions based on DNA sequences according to claim 2, characterized in that, The step of using the second granularity feature extraction module to perform multi-scale feature extraction on the multiple sets of k-mer processed promoter sequence fragments to obtain the first multi-scale feature, and to perform multi-scale feature extraction on the multiple sets of k-mer processed enhancer sequence fragments to obtain the second multi-scale feature, includes: By using multiple convolutional networks of different scales, feature extraction is performed on each group of promoter sequence fragments and each group of enhancer sequence fragments after k-mer processing, resulting in multiple promoter convolutional features and multiple enhancer convolutional features. Batch normalization and ReLU activation are performed on the multiple promoter convolutional features and the multiple enhancer convolutional features to obtain optimized promoter convolutional features and optimized enhancer convolutional features; Using a convolution kernel of size 1, both the optimized promoter convolutional feature and the optimized enhancer convolutional feature are subjected to secondary convolution to obtain the first multi-scale feature and the second multi-scale feature.

5. The method for predicting promoter-enhancer interactions based on DNA sequences according to claim 1, characterized in that, The method utilizes a gating mechanism to adaptively fuse the global representation of the promoter, the global representation of the enhancer, the first multi-scale feature, and the second multi-scale feature to obtain promoter fusion features and enhancer fusion features, including: Obtain the first gating weight, the second gating weight, and the third gating weight; The first gate weight and the global representation of the promoter are processed to obtain the first gate data; the second gate weight and the first multi-scale feature are processed to obtain the second gate data; and the third gate weight, the global representation of the promoter, and the first multi-scale feature are processed using the sigmoid activation function to obtain the third gate data. The first gate data, the second gate data, and the third gate data are weighted to obtain the promoter fusion feature; The first gate weight and the global representation of the enhancer are processed to obtain the fourth gate data; the second gate weight and the second multi-scale feature are processed to obtain the fifth gate data; and the third gate weight, the global representation of the enhancer and the second multi-scale feature are processed using the sigmoid activation function to obtain the sixth gate data. The fourth, fifth, and sixth gate data are weighted to obtain the enhancer fusion feature.

6. The method for predicting promoter-enhancer interactions based on DNA sequences according to claim 5, characterized in that, The first gating weight, the second gating weight, and the third gating weight are obtained by calculation using a deep learning network.

7. The method for predicting promoter-enhancer interactions based on DNA sequences according to claim 1, characterized in that, The method involves using a fixed-length window to truncate the acquired multiple sets of promoter sequences and multiple sets of enhancement sequences, resulting in multiple sets of promoter sequence fragments and multiple sets of enhancement sequence fragments, including: Using the window, each of the multiple sets of promoter sequences is slidably truncated to obtain the multiple sets of promoter sequence fragments; Using the window, each of the multiple sets of enhanced sub-sequences is trunculated to obtain the multiple sets of enhanced sub-sequence fragments.

8. The method for predicting promoter-enhancer interactions based on DNA sequences according to claim 4, characterized in that, The multiple convolutional networks with different scales include: convolutional kernels of size 3, size 5, and size 7 that are processed in parallel.

Citation Information

Patent Citations

  • Two-phase gating attention time sequence classification method and system based on two-way jump storage

    CN117828407A

  • Multi-source remote sensing optical image registration and fusion method based on space deformation field

    CN118941600A