Promoter-enhancer interaction prediction method based on DNA sequence
By adopting a DNA sequence-based method in promoter-enhancer interaction prediction, combining feature adaptive fusion and cross-attention calculation, the problem of insufficient prediction accuracy in the prior art is solved, and higher prediction accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510224961.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-27
AI Technical Summary
The prior art has insufficient feature extraction, feature fusion and prediction accuracy in the prediction of promoter-enhancer interactions, resulting in low prediction accuracy and reliability.
Using a prediction method based on DNA sequences, promoter and enhancer sequences are intercepted through a fixed-length window, and after k-mer processing, the promoter-enhancer prediction model is used for global feature extraction and multi-scale feature extraction, combining feature adaptive fusion and cross-attention calculation, an effective interaction modeling mechanism is established.
It significantly improves the model's prediction ability of promoter-enhancer interactions, and improves prediction accuracy and reliability.
Smart Images

Figure CN120072055A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics technology, and particularly relates to a method for predicting promoter-enhancer interactions based on DNA sequences. Background Art
[0002] An enhancer is an important non-coding regulatory element in the genome, which can regulate gene expression over a long distance by interacting with a promoter. Enhancer-promoter interaction (EPI) is a key mechanism in the gene expression regulatory network and is of great significance for cell differentiation, organ development, and disease occurrence.
[0003] The enhancer activates the promoter of its target gene through specific transcription factor binding sites, promoting the binding of RNA polymerase and the initiation of transcription. This interaction between the enhancer and the promoter is not limited to neighboring genes within the linear genomic distance, but can also act on genes at a long distance through three-dimensional genomic conformation. The function of EPI is crucial for the development process of cells, the morphogenesis of embryos, and the regulation of the immune system. Research shows that abnormal enhancer-promoter interactions may lead to various complex diseases, including cancer, autoimmune diseases, and neurodegenerative diseases. Therefore, predicting enhancer-promoter interactions can help identify key regulatory elements related to diseases and reveal the potential mechanisms of abnormal gene expression. For example, by predicting enhancers related to cancer driver genes, the transcriptional regulatory mechanism of tumorigenesis can be analyzed. EPI prediction can also provide potential drug targets, especially targets with inactivated or abnormally enhanced enhancers, providing a basis for personalized treatment. By reconstructing the genomic regulatory network, the understanding of complex phenotypes and polygenic traits can be improved, promoting the development of functional genomics.
[0004] Although the biological importance of enhancer-promoter interactions has been widely recognized, accurately predicting the interaction between enhancers and promoters still faces many challenges. Currently, there are various promoter-enhancer interaction prediction models, but these models still have obvious deficiencies in feature extraction, feature fusion, and prediction accuracy, seriously reducing the precision and reliability of EPI prediction. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present invention provides a method for predicting promoter-enhancer interactions based on DNA sequences. The technical problems to be solved by the present invention are achieved through the following technical solutions:
[0006] The present invention provides a method for predicting promoter-enhancer interactions based on DNA sequences, including:
[0007] Using a fixed - length window, intercept the obtained multiple groups of promoter sequences and multiple groups of enhancer sequences respectively to obtain multiple groups of promoter sequence fragments and multiple groups of enhancer sequence fragments; perform k - mer processing on each group of promoter sequence fragments and each group of enhancer sequence fragments respectively to obtain multiple groups of k - mer processed promoter sequence fragments and multiple groups of k - mer processed enhancer sequence fragments; based on the promoter - enhancer prediction model, perform global feature extraction and multi - scale feature extraction on the multiple groups of k - mer processed promoter sequence fragments respectively to obtain promoter global representations and first multi - scale features, and perform global feature extraction and multi - scale feature extraction on the multiple groups of k - mer processed enhancer sequence fragments respectively to obtain enhancer global representations and second multi - scale features; perform feature adaptive fusion and cross - attention calculation on the promoter global representation, the enhancer global representation, the first multi - scale feature, and the second multi - scale feature in sequence to obtain cross - attention results; calculate a prediction result based on the cross - attention results.
[0008] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0009] Aiming at the problem that there are still obvious deficiencies in feature extraction, feature fusion, and prediction accuracy of existing multiple promoter - enhancer interaction prediction models, which seriously reduce the accuracy and reliability of EPI prediction, the present invention provides a method for predicting promoter - enhancer interaction based on DNA sequences. By using a promoter - enhancer prediction model, this method performs global feature extraction and multi - scale feature extraction on multiple groups of k - mer processed promoter sequence fragments respectively to capture information in DNA sequences at multiple levels and scales, and through the way of feature adaptive fusion, adaptively adjusts the contribution degrees of features at different scales to fully consider the importance of different features and fuse feature information at different scales and complexities. Subsequently, cross - attention calculation is used to establish an effective interaction modeling mechanism to accurately capture the interaction pattern between promoters and enhancers. Finally, a prediction result is calculated using the cross - attention results, significantly improving the model's prediction ability for promoter - enhancer interaction and effectively enhancing the prediction accuracy and reliability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a schematic flowchart of a method for predicting promoter - enhancer interaction based on DNA sequences provided by an embodiment of the present invention;
[0011] Figure 2 is an application example diagram of the method for predicting promoter - enhancer interaction based on DNA sequences provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0012] The present invention will be further described in detail below in conjunction with specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0013] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means two or more, unless otherwise specifically defined.
[0014] In the description of this specification, the description with reference to terms such as "an embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0015] Although the present invention has been described in conjunction with various embodiments herein, however, in the process of implementing the claimed invention, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality of cases. A single processor or other unit can implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0016] Now, in conjunction with the accompanying drawings, a method for predicting promoter-enhancer interactions based on DNA sequences provided by the present invention will be described in detail.
[0017] Figure 1 It is a schematic flowchart of a method for predicting promoter-enhancer interactions based on DNA sequences provided by an embodiment of the present invention. As Figure 1 shown, the method includes steps 110-150. Specifically:
[0018] Step 110: Use a window with a fixed length to respectively intercept multiple groups of promoter sequences and multiple groups of enhancer sequences obtained, to obtain multiple groups of promoter sequence fragments and multiple groups of enhancer sequence fragments.
[0019] Here, step 110 specifically includes: using a window to perform sliding truncation on each group of promoter sequences in multiple groups of promoter sequences to obtain multiple groups of promoter sequence fragments; using a window to perform sliding truncation on each group of enhancer sequences in multiple groups of enhancer sequences to obtain multiple groups of enhancer sequence fragments.
[0020] Exemplarily, multiple groups of promoter sequences P and multiple groups of enhancer sequences E belong to a verified promoter-enhancer interaction dataset in the publicly available human genome. And, a group of promoter sequences P and a group of enhancer sequences E form a pair of promoter-enhancer sequences. The number of promoter sequences P and enhancer sequences E is the same, and the number of fragments obtained by truncating both is also the same. Here, the length of the window can be 2 kbp or 3 kbp. It should be understood that the present invention does not limit the size of the window length and can be adjusted according to actual situations.
[0021] Step 120: Perform k-mer processing on each group of promoter sequence fragments and each group of enhancer sequence fragments respectively to obtain multiple groups of k-mer processed promoter sequence fragments and multiple groups of k-mer processed enhancer sequence fragments.
[0022] Here, the core of k-mer processing is to map each small segment (with length k) in the DNA sequence to a numerical representation using the sliding window technique. Taking k = 6 as an example, for the sequence ACGTAGCTG, k-mer processing will convert it into four segments [ACGTAG, CGTAGC, GTAGCT, TAGCTG], and each segment is converted into a vector with a fixed length.
[0023] Here, through k-mer processing, the fixed-length sequence fragments are divided again, enabling the promoter sequences P and enhancer sequences E to be used as input data in subsequent deep learning models. Generally, the value of k ranges from 5 to 10. A shorter k value helps to capture local sequence features, and a longer k value can capture more global information. Exemplarily, assuming the fixed-length promoter sequence P is represented as Then, the k-mer processed promoter sequence fragments can be represented as: Similarly, the fixed-length enhancer sequence E is represented as Then, the k-mer processed promoter sequence fragments can be represented as:
[0024] Step 130: Based on the promoter-enhancer prediction model, perform global feature extraction and multi-scale feature extraction on multiple groups of promoter sequence fragments after k-mer processing respectively to obtain the promoter global representation and the first multi-scale features, and perform global feature extraction and multi-scale feature extraction on multiple groups of enhancer sequence fragments after k-mer processing respectively to obtain the enhancer global representation and the second multi-scale features.
[0025] Specifically, the promoter-enhancer prediction model includes a first granularity feature extraction module and a second granularity feature extraction module; Step 130 includes: (1) Using the first granularity feature extraction module, perform global feature extraction on multiple groups of promoter sequence fragments after k-mer processing to obtain the promoter global representation, and perform global feature extraction on multiple groups of enhancer sequence fragments after k-mer processing to obtain the enhancer global representation; (2) Using the second granularity feature extraction module, perform multi-scale feature extraction on multiple groups of promoter sequence fragments after k-mer processing to obtain the first multi-scale features, and perform multi-scale feature extraction on multiple groups of enhancer sequence fragments after k-mer processing to obtain the second multi-scale features.
[0026] It should be noted that the granularity of the first granularity feature extraction module is greater than that of the second granularity feature extraction module. In other words, the first granularity feature extraction module can be understood as a coarse-grained feature extraction module, and the second granularity feature extraction module can be understood as a fine-grained feature extraction module.
[0027] Here, step (1) further includes: Inputting each group of promoter sequence fragments after k-mer processing and each group of enhancer sequence fragments after k-mer processing into the pre-trained Nucleotide Transformers model respectively to obtain the corresponding first context-aware representation sequence and the second context-aware representation sequence; Performing global max pooling and global average pooling on each group of the first context-aware representation sequences to obtain the corresponding first max pooling data and first average pooling data, and performing global max pooling and global average pooling on each group of the second context-aware representation sequences to obtain the corresponding second max pooling data and second average pooling data; Concatenating the first max pooling data and the first average pooling data corresponding to each group of promoter sequence fragments after k-mer processing to obtain the promoter global representation; Concatenating the second max pooling data and the second average pooling data corresponding to each group of enhancer sequence fragments after k-mer processing to obtain the enhancer global representation.
[0028] Here, context-aware sequence representation means that when performing sequence modeling, the model can understand and utilize the context information around each element in the sequence, so as to more accurately predict or generate the sequence content, and the NucleotideTransformers model refers to the transformer model proposed based on amino acid gene sequences. Here, the pre-trained Nucleotide Transformers model uses the Rotary Position Encoding (RoPE) algorithm to process the promoter sequence fragments after processing each group of k-mers and the enhancer sequence fragments after processing each group of k-mers respectively. The RoPE algorithm takes advantage of the frequency (sinusoidal) distribution in position encoding and combines the traditional sine wave method. Specifically, for each position, RoPE calculates its frequency range and then uses the frequency value for complex rotation. The calculation of frequency is related to the position, similar to the traditional sine wave encoding, but in RoPE, the position vector is rotated in each dimension, so that the relative position information can be captured. During the rotation process, each token rotates by different angles in different dimensions, representing the relative position of that position. RoPE actually combines the rotation characteristics of sine and cosine functions, enabling it to better represent relative position information. Specifically, RoPE calculates a set of rotation matrices for each position and then uses them to transform the input token embeddings. This transformation helps the model learn the relationships between adjacent tokens and the global relative position relationships.
[0029] Moreover, compared with the traditional position encoding that represents the uniqueness of each position through fixed sine and cosine wave functions, giving an absolute position identifier to each position, and this encoding is fixed and does not change with the training data, RoPE does not directly give the absolute value of each position, but captures the relative position relationship through rotational encoding. This approach makes RoPE have better flexibility and scalability and can maintain good performance on inputs of different lengths.
[0030] In one possible implementation, the first context-aware representation sequence H P can be expressed as: H P = Nucleotide Transformers(X P ), and, the second context-aware representation sequence H E can be expressed as: H E = Nucleotide Transformers(X E ).
[0031] It should be noted that there is no order of precedence for the processing of global max pooling and global average pooling. Among them, global max pooling extracts the maximum value of each feature dimension in the sequence, while global average pooling performs an averaging operation on each feature dimension. The information extracted by max pooling and average pooling from the sequence is complementary. Max pooling focuses on significant features, while average pooling focuses on the overall features of the sequence. By concatenating the results of both, the strongest signals and global information can be retained simultaneously, thus obtaining a richer and more comprehensive feature representation.
[0032] In a possible implementation, the global promoter representation P global can be expressed as: P global = Concat(MaxPool(H P ), AvgPool(H P )); where MaxPool(·) is the global max pooling process, AvgPool(·) is the global average pooling, and GlobalFeatureConcat is the global feature concatenation.
[0033] In a possible implementation, the global enhancer representation can be expressed as: E global = Concat(MaxPool(H E ), AvgPool(H E ));
[0034] Here, step (2) specifically includes: using multiple convolutional networks with different scales to separately extract features from the promoter sequence fragments processed by each group of k-mers and the enhancer sequence fragments processed by each group of k-mers, obtaining multiple promoter convolutional features and multiple enhancer convolutional features; performing batch normalization processing and ReLU activation processing on both the multiple promoter convolutional features and the multiple enhancer convolutional features to obtain optimized promoter convolutional features and optimized enhancer convolutional features; using a convolutional kernel of size 1 to perform secondary convolution on both the optimized promoter convolutional features and the optimized enhancer convolutional features to obtain the first multi-scale features and the second multi-scale features.
[0035] Exemplarily, the multiple convolutional networks with different scales include: convolutional kernels of size 3, size 5, and size 7 that are processed in parallel. Multi-scale local features are extracted by using convolutional layers with different convolutional kernel sizes. Here, convolutional kernels of size 3, size 5, and size 7 are used to separately extract features from the promoter sequence fragments processed by each group of k-mers and the enhancer sequence fragments processed by each group of k-mers. After completing the feature extraction, batch normalization processing and ReLU activation processing are used to enhance the non-linear expression ability and reduce the internal covariance shift.
[0036] In a possible implementation, the calculation process of each convolutional kernel can be expressed as: ConV(x) k = ReLU(W k * x + b k ), where W k is the convolutional kernel, b k is the bias term, whose parameters are automatically changed following model training, and * is the convolution operation.
[0037] Step 140: Perform feature adaptive fusion and cross-attention calculation on the promoter global representation, enhancer global representation, first multi-scale feature, and second multi-scale feature in sequence to obtain the cross-attention result.
[0038] Here, step 140 specifically includes: (a) Using a gating mechanism, perform feature adaptive fusion on the promoter global representation, enhancer global representation, first multi-scale feature, and second multi-scale feature to obtain the promoter fusion feature and enhancer fusion feature; (b) Using the promoter fusion feature and enhancer fusion feature, construct an attention key-value pair for characterizing the similarity between the promoter global representation and the enhancer global representation; (c) Using the attention key-value pair, calculate the cross-attention result.
[0039] Among them, step (a) specifically includes: obtaining the first gating weight, the second gating weight, and the third gating weight; performing arithmetic processing on the first gating weight and the promoter global representation to obtain the first gate data, performing arithmetic processing on the second gating weight and the first multi-scale feature to obtain the second gate data, and using the sigmoid activation function to perform arithmetic processing on the third gating weight, the promoter global representation, and the first multi-scale feature to obtain the third gate data; performing weighted processing on the first gate data, the second gate data, and the third gate data to obtain the promoter fusion feature; performing arithmetic processing on the first gating weight and the enhancer global representation to obtain the fourth gate data, performing arithmetic processing on the second gating weight and the second multi-scale feature to obtain the fifth gate data, and using the sigmoid activation function to perform arithmetic processing on the third gating weight, the enhancer global representation, and the second multi-scale feature to obtain the sixth gate data; performing weighted processing on the fourth gate data, the fifth gate data, and the sixth gate data to obtain the enhancer fusion feature. It should be noted that the first gating weight, the second gating weight, and the third gating weight are obtained by calculating using a deep learning network.
[0040] In a possible implementation, the promoter fusion feature P fused can be expressed as:
[0041] P fused = z · h g + (1 - z) · h l ;
[0042] h g = tanh(W g ·P global );
[0043] h l = tanh(W l ·P local );
[0044] z = σ(W z ·[P global , P local );
[0045] Among them, P local refers to the first multi-scale feature, W g is the first gating weight, W l is the second gating weight, W z is the third gating weight, and σ is the sigmoid activation function.
[0046] In a possible implementation, the enhancer fusion feature E fused can be expressed as:
[0047] E fused = z · h g + (1 - z) · h l ;
[0048] h g = tanh(W g · E global );
[0049] h l = tanh(W l · E local );
[0050] z = σ(W z · [E global , E local );
[0051] Among them, E local refers to the second multi-scale feature.
[0052] Here, the attention key-value pairs in step (b) include: the query matrix Q P responsible for mapping the promoter feature to the query space, the key matrix K E , and the value matrix V E . Among them, the expression of the query matrix Q P is: Q P = P fused W Q , and the expression of the key matrix K E is: K E = Efused W K , the value matrix V E has the expression: V E = E fused W V .
[0053] Here, the cross-attention result in step (c) is expressed as: CrossAtten(P, E) = A·V E , where A is the attention weight matrix, representing the similarity between promoter and enhancer features, and its expression is: Here, D refers to the attention scaling parameter.
[0054] Step 150: Calculate the prediction result based on the cross-attention result.
[0055] Here, step 150 specifically includes: inputting the cross-attention result into a multi-layer perceptron to calculate and obtain the prediction result. Among them, the multi-layer perceptron includes several fully connected layers. The multi-layer perceptron (MLP) is a feedforward artificial neural network composed of multiple neurons (nerve nodes). These neurons are arranged in a hierarchical structure, including an input layer, a hidden layer, and an output layer. Neurons between layers are connected by weights, and information propagates forward from the input layer to the output layer in sequence without feedback connections. After each layer in the MLP is operated, to avoid overfitting, the output value of each layer is processed using layer normalization (LayerNorm) and regularization (dropout).
[0056] Here, the multi-layer perceptron uses the binary cross-entropy loss function to constrain the output value, and its expression formula is: Loss = -∑(y true *log(y pred ) + (1 - y true )*log(1 - y pred ))). Furthermore, the calculation formula for the output value Y of the multi-layer perceptron is Y = MLP(CrossAtten(P, E)), and the prediction result can be expressed as: y pred = σ(Y).
[0057] Figure 2 is an application example diagram of the promoter-enhancer interaction prediction method based on DNA sequence provided by the embodiments of the present invention. As Figure 2As shown, after multiple pairs of promoter-enhancer sequences are normalized (i.e., intercepted using a fixed-length window) and k-mer processed, they are respectively input into the first granularity feature extraction module and the second granularity feature extraction module in the promoter-enhancer prediction model. The overall feature information of the sequence is extracted by the first granularity feature extraction module to obtain the promoter-enhancer global representation, and the local feature information of the sequence is extracted from multiple levels and multiple scales by the second granularity feature module to obtain multiple groups of local features. Subsequently, a gated unit is used to adaptively fuse the input multiple feature information based on actual requirements and the importance of each feature, and a cross-attention unit is used to perform interactive calculations on the fused feature information. Then, a prediction unit predicts the result of the interactive calculation to obtain the final prediction result.
[0058] Aiming at the problem that there are still obvious deficiencies in the existing multiple promoter-enhancer interaction prediction models in terms of feature extraction, feature fusion, and prediction accuracy, which seriously reduce the accuracy and reliability of EPI prediction, the present invention provides a method for predicting the interaction between a promoter and an enhancer based on a DNA sequence. This method uses a promoter-enhancer prediction model to respectively perform global feature extraction and multi-scale feature extraction on multiple groups of promoter sequence fragments after k-mer processing, capture the information in the DNA sequence from multiple levels and multiple scales, and adaptively adjust the contribution degrees of features at different scales through feature adaptive fusion to fully consider the importance of different features and fuse feature information at different scales and complexities. Subsequently, cross-attention calculation is used to establish an effective interaction modeling mechanism to accurately capture the interaction mode between the promoter and the enhancer. Finally, the prediction result is calculated using the cross-attention result, significantly improving the prediction ability of the model for promoter-enhancer interaction and effectively improving the accuracy and reliability of the model prediction.
[0059] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should all be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for predicting promoter-enhancer interactions based on DNA sequences, characterized in that: include: Using a window of fixed length, the obtained multiple groups of promoter sequences and multiple groups of enhancer sequences are respectively intercepted to obtain multiple groups of promoter sequence fragments and multiple groups of enhancer sequence fragments; Performing k-mer processing on each group of promoter sequence fragments and each group of enhancer sequence fragments respectively, to obtain multiple groups of promoter sequence fragments after k-mer processing, and multiple groups of enhancer sequence fragments after k-mer processing; Based on the promoter-enhancer prediction model, performing global feature extraction and multi-scale feature extraction on the multiple groups of promoter sequence fragments after k-mer processing, respectively, to obtain promoter global representation and first multi-scale features, and performing global feature extraction and multi-scale feature extraction on the multiple groups of enhancer sequence fragments after k-mer processing, respectively, to obtain enhancer global representation and second multi-scale features; performing feature adaptive fusion and cross-attention calculation on the promoter global representation, the enhancer global representation, the first multi-scale feature, and the second multi-scale feature in sequence to obtain a cross-attention result; Based on the cross-attention results, a prediction result is calculated.
2. The promoter-enhancer interaction prediction method based on DNA sequence according to claim 1, characterized in that: The promoter-enhancer prediction model includes a first granularity feature extraction module and a second granularity feature extraction module; Based on the promoter-enhancer prediction model, the global feature extraction and multi-scale feature extraction are respectively performed on the multiple groups of promoter sequence fragments after k-mer processing to obtain the promoter global representation and the first multi-scale feature, and the global feature extraction and multi-scale feature extraction are respectively performed on the multiple groups of enhancer sequence fragments after k-mer processing to obtain the enhancer global representation and the second multi-scale feature, including: Using the first granularity feature extraction module, performing global feature extraction on the multiple groups of promoter sequence fragments after k-mer processing to obtain the promoter global representation, and performing global feature extraction on the multiple groups of enhancer sequence fragments after k-mer processing to obtain the enhancer global representation; The second granularity feature extraction module is used to perform multi-scale feature extraction on the multiple groups of promoter sequence fragments after k-mer processing to obtain the first multi-scale features, and multi-scale feature extraction is performed on the multiple groups of enhancer sequence fragments after k-mer processing to obtain the second multi-scale features.
3. The method for predicting promoter-enhancer interactions based on DNA sequences according to claim 2, characterized in that: The method of using the first granularity feature extraction module to perform global feature extraction on the multiple groups of promoter sequence fragments processed by k-mer to obtain the promoter global representation, and performing global feature extraction on the multiple groups of enhancer sequence fragments processed by k-mer to obtain the enhancer global representation, includes: Input each group of k-mer processed promoter sequence fragments and each group of k-mer processed enhancer sequence fragments into the pre-trained Nucleotide Transformers model to obtain the corresponding first context-aware representation sequence and second context-aware representation sequence; Performing global maximum pooling and global average pooling on each group of first context-aware representation sequences to obtain corresponding first maximum pooling data and first average pooling data, and performing global maximum pooling and global average pooling on each group of second context-aware representation sequences to obtain corresponding second maximum pooling data and second average pooling data; splicing the first maximum pooling data and the first average pooling data respectively corresponding to the multiple groups of promoter sequence fragments after k-mer processing to obtain the promoter global representation; The second maximum pooling data and the second average pooling data corresponding to the multiple groups of enhancer sequence fragments after k-mer processing are spliced to obtain the global representation of the enhancer.
4. The method for predicting promoter-enhancer interactions based on DNA sequences according to claim 2, characterized in that: The method of using the second granularity feature extraction module to perform multi-scale feature extraction on the multiple groups of promoter sequence fragments processed by k-mer to obtain the first multi-scale feature, and performing multi-scale feature extraction on the multiple groups of enhancer sequence fragments processed by k-mer to obtain the second multi-scale feature, includes: Using multiple convolutional networks of different scales, feature extraction is performed on each group of promoter sequence fragments processed by k-mer and each group of enhancer sequence fragments processed by k-mer, to obtain multiple promoter convolution features and multiple enhancer convolution features; Performing batch normalization processing and ReLU activation processing on the multiple promoter convolution features and the multiple enhancer convolution features to obtain optimized promoter convolution features and optimized enhancer convolution features; Using a convolution kernel with a size of 1, a secondary convolution is performed on the optimized promoter convolution feature and the optimized enhancer convolution feature to obtain the first multi-scale feature and the second multi-scale feature.
5. The method for predicting promoter-enhancer interactions based on DNA sequences according to claim 1, characterized in that: The step of sequentially performing feature adaptive fusion and cross-attention calculation on the promoter global representation, the enhancer global representation, the first multi-scale feature, and the second multi-scale feature to obtain a cross-attention result includes: Using a gating mechanism, adaptively fusing the promoter global representation, the enhancer global representation, the first multi-scale feature, and the second multi-scale feature to obtain a promoter fusion feature and an enhancer fusion feature; Using the promoter fusion feature and the enhancer fusion feature, constructing an attention key-value pair for characterizing the similarity between the promoter global representation and the enhancer global representation; The cross-attention result is calculated using the attention key-value pairs.
6. The method for predicting promoter-enhancer interactions based on DNA sequences according to claim 5, characterized in that: The method of using a gating mechanism to adaptively fuse the promoter global representation, the enhancer global representation, the first multi-scale feature, and the second multi-scale feature to obtain a promoter fusion feature and an enhancer fusion feature includes: Obtaining a first gating weight, a second gating weight, and a third gating weight; Performing operation processing on the first gating weight and the promoter global representation to obtain first gate data, performing operation processing on the second gating weight and the first multi-scale feature to obtain second gate data, and performing operation processing on the third gating weight, the promoter global representation and the first multi-scale feature using a sigmoid activation function to obtain third gate data; Performing weighted processing on the first gate data, the second gate data and the third gate data to obtain the promoter fusion feature; The first gating weight and the enhancer global representation are processed by operation to obtain fourth gate data, the second gating weight and the second multi-scale feature are processed by operation to obtain fifth gate data, and the third gating weight, the enhancer global representation and the second multi-scale feature are processed by operation using a sigmoid activation function to obtain sixth gate data; The fourth gate data, the fifth gate data and the sixth gate data are weighted to obtain the enhancer fusion feature.
7. The method for predicting promoter-enhancer interactions based on DNA sequences according to claim 6, characterized in that: The first gating weight, the second gating weight and the third gating weight are calculated using a deep learning network.
8. The method for predicting promoter-enhancer interactions based on DNA sequences according to claim 1, characterized in that: The calculation based on the cross attention result to obtain the prediction result includes: The cross attention result is input into a multi-layer perceptron to calculate and obtain the prediction result, wherein the multi-layer perceptron includes a plurality of fully connected layers.
9. The method for predicting promoter-enhancer interactions based on DNA sequences according to claim 1, characterized in that: The method uses a fixed-length window to intercept the acquired multiple groups of promoter sequences and multiple groups of enhancer sequences respectively to obtain multiple groups of promoter sequence fragments and multiple groups of enhancer sequence fragments, including: Using the window, slidingly intercepting each group of promoter sequences in the multiple groups of promoter sequences to obtain the multiple groups of promoter sequence fragments; The window is used to perform sliding interception on each group of enhancer sequences in the multiple groups of enhancer sequences to obtain the multiple groups of enhancer sequence fragments.
10. The method for predicting promoter-enhancer interactions based on DNA sequences according to claim 4, characterized in that: The multiple convolutional networks of different scales include: a convolution kernel of size 3, a convolution kernel of size 5, and a convolution kernel of size 7 processed in parallel.
Citation Information
Patent Citations
Multi-head attention mechanism-based enhancer-promoter interaction prediction model construction method
CN116312748A
Two-phase gating attention time sequence classification method and system based on two-way jump storage
CN117828407A
Multi-source remote sensing optical image registration and fusion method based on space deformation field
CN118941600A
Controllable video generation method and system based on multi-modal fusion
CN119091362A
Chromosome neighborhood structures and methods relating thereto
US20190005191A1