A Transformer model for predicting enhancer-promoter interactions based on transcription factor cross-attention

By combining a transformer model based on transcription factor cross-attention with ChIP-seq data and the Cross-Attention mechanism, the problem of insufficient biological interpretability in existing methods was solved, and high-precision enhancer-promoter interaction prediction and regulatory relationship simulation were achieved.

CN120183487BActive Publication Date: 2025-09-12YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510655862.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-09-12
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

Existing enhancer-promoter interaction prediction methods ignore the interaction patterns of transcription factors, have limited biological interpretability, and find it difficult to achieve high-precision predictions.

Method used

A transformer model based on cross-attention prediction of transcription factors is used, combined with the spatial order and binding strength information extracted from ChIP-seq peaks to construct a transcription factor sequence feature representation. The Cross-Attention mechanism in the Transformer architecture is introduced to model the long-range regulatory relationship between enhancers and promoters.

Benefits of technology

It achieves high-precision prediction of enhancer-promoter interactions, has stronger biological explanatory power, can simulate complex cross-regional regulatory interactions, and support subsequent biological function analysis and regulatory mechanism exploration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183487B_ABST
    Figure CN120183487B_ABST
Patent Text Reader

Abstract

This paper provides a Transformer model for predicting enhancer-promoter interactions based on transcription factor cross-attention. This model addresses the limited interpretability of existing enhancer-promoter interaction prediction methods. The model includes an input layer connected to a sine-cosine positional encoding module via a convolutional feature extraction module, which in turn is connected to a multi-layer feedforward network module via a cross-attention encoder layer. This model offers advantages such as good interpretability and high prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical care informatics, and in particular relates to a transformer model for predicting enhancer-promoter interactions based on transcription factor cross-attention. Background Art

[0002] Currently, a variety of methods are used to predict enhancer-promoter interactions, including machine learning-based prediction models (TargetFinder, kmer-SVM, EP2vec) and deep learning-based models (EPIANN, SPEID, EPI-DLMH, and EPI-trans). Existing enhancer-promoter interaction prediction methods primarily rely on DNA sequence features and epigenetic modifications. These methods ignore the transcription factor interactions that directly mediate enhancer-promoter interactions, and their biological interpretability remains limited. Therefore, a framework that can explicitly model and interpret the regulatory interactions between transcription factors is urgently needed to improve the model's ability to analyze enhancer-promoter interactions.

[0003] This scheme proposes a transformer model for predicting enhancer-promoter interactions based on transcription factor cross-attention. It combines the spatial order and binding strength information extracted from ChIP-seq peaks to construct a transcription factor sequence feature representation, and introduces the Cross-Attention mechanism in the Transformer architecture to effectively model the long-range regulatory relationship between enhancers and promoters, achieving high-precision prediction of enhancer-promoter interactions. Summary of the Invention

[0004] The purpose of the present invention is to address the above problems and provide a transformer model with a reasonable design and good interpretability for predicting enhancer-promoter interactions based on transcription factor cross-attention.

[0005] Another object of the present invention is to address the above-mentioned problems and provide a method for predicting enhancer-promoter interactions based on transcription factor cross-attention with high prediction accuracy.

[0006] To achieve the above objectives, the present invention adopts the following technical solutions: a transformer model based on transcription factor cross-attention prediction of enhancer-promoter interaction, including an input layer, the input layer is connected to a sine-cosine position encoding module through a convolutional feature extraction module, and the sine-cosine position encoding module is connected to a multi-layer feedforward network module through a cross-attention encoder layer.

[0007] A method for predicting enhancer-promoter interactions based on transcription factor cross-attention prediction, comprising the following steps:

[0008] S1: Data acquisition and preprocessing, building a data foundation for enhancer-promoter interaction prediction;

[0009] S2: Standardization of enhancer and promoter regions, unified region definition and coordinate processing;

[0010] S3: Key transcription factor feature encoding and input tensor generation, converting key transcription factor binding information into structured input;

[0011] S4: Model architecture design, building a transformer model that integrates local features and global interactions;

[0012] S5: Model training and evaluation, optimization measurement and robustness verification;

[0013] S6: We use the cross-attention mechanism within the transformer model to perform interpretable analysis of transcription factor interactions between enhancer and promoter regions.

[0014] In the above-mentioned enhancer-promoter mutual prediction method based on transcription factor cross-attention prediction, step S1 includes the following steps:

[0015] S11: Download ChIP-seq data of 26 key transcription factors in GM12878 and K562 cell lines;

[0016] S12: Obtain files containing enhancer and promoter pairing information, transcription start site information, and training set files for GM12878 and K562 cell types.

[0017] In the above-mentioned enhancer-promoter mutual prediction method based on transcription factor cross-attention prediction, step S2 includes the following steps:

[0018] S21: Enhancer processing, extending 3 kb to both sides of the enhancer center to generate a 6 kb normalized region;

[0019] S22: promoter processing, integrating the transcription start site information and defining the transcription factor site ± 2.5 kb as the promoter region;

[0020] S23: Data screening: retain valid samples with strand information matching within the enhancer-promoter pairing region, and remove samples with many-to-many matching and no strand matching.

[0021] In the above-mentioned enhancer-promoter mutual prediction method based on transcription factor cross-attention prediction, step S3 includes the following steps:

[0022] S31: Parse the ChIP-seq file of transcription factors to extract the binding peak position and binding intensity of transcription factors;

[0023] S32: Assign a unique character code to each transcription factor, build a unified coding dictionary and store it as a JSON file;

[0024] S33: For each enhancer or promoter region, extract all overlapping transcription factor binding sites within its coordinate range, generate transcription factor coding sequences in positional order, and record the corresponding binding strength numerical values. If it is a negative strand promoter, its sequence is reversed;

[0025] S34: One-hot encode the transcription factor sequence and add an additional column to represent the binding strength of the transcription factor at that position to generate a tensor representation;

[0026] S35: Set the maximum length for enhancer and promoter sequences respectively, fill the insufficient part with zero vector, truncate the excessive part, and output the standardized model input tensor.

[0027] In the above-mentioned enhancer-promoter mutual prediction method based on transcription factor cross-attention prediction, step S4 includes the following steps:

[0028] S41: Input enhancer and promoter sequence tensors, unifying sequence length and site feature dimensions;

[0029] S42: Extracting local sequential patterns using different convolution kernels, followed by ReLU activation, batch normalization, and max pooling;

[0030] S43: Perform sine and cosine position encoding and feature fusion;

[0031] S44: Perform bidirectional attention calculation, perform nonlinear mapping through the feedforward network, followed by residual connection and layer normalization.

[0032] In the above-mentioned method for predicting enhancer-promoter mutual prediction based on transcription factor cross-attention, step S4 inputs the feature maps of the enhancer and promoter after passing through the cross-attention encoder layer into the global average pooling layer, compresses the sequence features of different lengths into vector representations of fixed dimensions, splices the representation vectors of the two regions into an overall joint feature vector, and performs nonlinear transformation and feature fusion through a fully connected neural network module with a Dropout mechanism. The fused features are passed through a Sigmoid-activated output layer to generate a probability value between 0 and 1. When the value is greater than the preset threshold, the model predicts that the pair has a potential long-range regulatory relationship.

[0033] In the above-mentioned transcription factor cross-attention-based enhancer-promoter mutual prediction method, step S4 stacks three layers of cross-attention encoder layers on each channel.

[0034] In the above-mentioned enhancer-promoter mutual prediction method based on transcription factor cross-attention prediction, step S5 uses a ten-fold cross-validation method to divide the data in the training stage, and uses weighted binary cross entropy combined with category weights to process unbalanced data for the loss function. The Adam optimizer is used for the optimizer, and the initial learning rate is set to 1e-3 and the parameters are adjusted dynamically.

[0035] In the above-mentioned enhancer-promoter mutual prediction method based on transcription factor cross-attention prediction, step S5 sets a multi-dimensional indicator system to evaluate model performance, including accuracy, area under the receiver operating characteristic curve, and area under the precision-recall curve.

[0036] Compared with existing technologies, the advantages of the present invention are: encoding the binding order and strength information of transcription factors in enhancer and promoter regions into transcription factor sequences, thereby directly mining gene regulation patterns from transcription factor binding events. Compared with methods that rely on DNA sequences or chromatin modifications, this method is closer to the actual transcriptional regulation process and has stronger biological explanatory power; the model uses a Cross-Attention structure based on the Transformer framework to perform bidirectional cross-attention modeling on enhancer and promoter features respectively, effectively simulating the complex cross-regional regulatory interactions between enhancers and promoters; the model attention output can be used to locate key TF combinations, and support subsequent biological function analysis and regulatory mechanism exploration by counting high-frequency interactions, constructing TF regulatory networks and performing community detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a workflow diagram of the present invention;

[0038] Figure 2 It is a model architecture diagram of the present invention;

[0039] Figure 3 This is a graph showing the results of a ten-fold cross validation of the K562 cell line of the present invention;

[0040] Figure 4 This is a graph showing the results of a ten-fold cross validation of the GM12878 cell line of the present invention;

[0041] Figure 5 This is a comparison result diagram of the down-sampling data of the K562 cell line of the present invention;

[0042] Figure 6 This is a comparison result diagram of the original data of the K562 cell line of the present invention;

[0043] Figure 7 This is a comparison result diagram of the down-sampling data of the GM12878 cell line of the present invention;

[0044] Figure 8 This is a comparison result diagram of the original data of the GM12878 cell line of the present invention;

[0045] Figure 9 is the K562 cell line enhancer-promoter sample attention heat map of the present invention;

[0046] Figure 10 is the GM12878 cell line enhancer-promoter sample attention heat map of the present invention;

[0047] Figure 11 is a high-frequency interaction map of transcription factors of the present invention;

[0048] Figure 12 is a frequency distribution diagram of TF in enhancer and promoter regions in the high attention region of the K562 cell line of the present invention;

[0049] Figure 13 is a frequency distribution diagram of TF in enhancer and promoter regions in the high attention region of the GM12878 cell line of the present invention;

[0050] Figure 14 is the TF interaction network diagram of the K562 cell line of the present invention;

[0051] Figure 15 This is the TF interaction network diagram of the GM12878 cell line of the present invention. DETAILED DESCRIPTION

[0052] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0053] like Figures 1-15 As shown, the present application provides a transformer model and mutual prediction method for predicting enhancer-promoter interactions based on transcription factor cross-attention. Based on dual-input sequence encoding, convolutional feature extraction and cross-attention mechanism, it identifies long-range regulatory relationships and combines the internal attention weights of the model to realize the interpretable analysis of the regulatory mechanism.

[0054] The model includes an input layer, which is connected to a sine-cosine position encoding module through a convolutional feature extraction module, and the sine-cosine position encoding module is connected to a cross-attention encoder layer and a multi-layer feedforward network module.

[0055] The forecasting method used includes the following steps:

[0056] S1: Data acquisition and preprocessing, building a data foundation for enhancer-promoter interaction prediction;

[0057] S2: Standardization of enhancer and promoter regions, unified region definition and coordinate processing;

[0058] S3: Key transcription factor feature encoding and input tensor generation, converting key transcription factor binding information into structured input;

[0059] S4: Model architecture design, building a transformer model that integrates local features and global interactions;

[0060] S5: Model training and evaluation, optimization measurement and robustness verification;

[0061] S6: We use the cross-attention mechanism within the transformer model to perform interpretable analysis of transcription factor interactions between enhancer and promoter regions.

[0062] Specifically, step S1 first downloaded ChIP-seq data (narrowPeak format files) for 26 transcription factors shared by the GM12878 and K562 cell lines and closely related to transcription from the ENCODE database. Then, the following files were downloaded from the TargetFinder project: the pairs.csv file containing enhancer and promoter pairing information, the gencode.v19.TSS.notlow.gff.gz file containing transcription start site (TSS) information from GENCODE v19 annotations, the BED format file TSS.bed containing transcription start site information, and the training set files K562train.csv and GM12878train.csv for K562 and GM12878 cell types from the EP2vec project. These training sets were downsampled at a 1:1 ratio.

[0063] In addition, when processing enhancer and promoter regions, step S2 first performs a centering operation on the input enhancer coordinates, and extends 3kb on both sides based on the center point to ensure that the region length is 6kb, so as to facilitate the subsequent unified processing of ChIP-seq peaks. The updated coordinate information will be used for subsequent matching and model encoding. The processing of promoter regions depends on two files: gencode.v19.TSS.notlow.gff and TSS.bed. By uniquely matching the chrom, start, and gene_id fields and integrating the strand information of the TSS region, a promoter region annotation with a clear direction is finally obtained. On this basis, the TSS site ±2.5kb is defined as the promoter region as the model input. For data screening and merging, valid samples with strand information matching in the enhancer-promoter pairing region are retained, while samples with many-to-many and no strand matching are removed to ensure the accuracy of downstream transcription factor mapping.

[0064] Furthermore, step S3 first parses multiple narrowPeak format TF ChIP-seq files to extract the summit position (i.e., binding peak apex) and signalValue (binding intensity) of each TF. Subsequently, a unique character code is assigned to each TF ("FOS":"A","ELK1":"B", "SMAD5":"C","BHLHE40":"D","NFATC3":"E","SP1":"F","YY1":"G","JUND":"H","NFIC":"I","TBP":"J","MYC":"K","CUX1":"L","EGR1":"M","SPI1":"N","ZNF384":"O","RFX5":" "P", "MAX": "Q", "MAFK": "R", "CEBPB": "S", "GABPA": "T", "POLR2A": "U", "CTCF": "V", "MXI1": "W", "NRF1": "X", "RAD21": "Y", "MAZ": "Z"), a unified encoding dictionary was constructed and stored as a JSON file to ensure consistency and traceability in subsequent data processing. For each enhancer or promoter region, all overlapping TF binding sites within its coordinate range were extracted. TF coding sequences (e.g., "ADHF") were generated in order of position, and a list of corresponding signalValue values ​​was recorded. For minus-strand promoters, the sequence was reversed. Based on this, the TF sequence was one-hot encoded, and an additional column was added to represent the signalValue of the TF at that position. The final tensor representation was generated with a shape of [region length, TF type + 1]. To ensure uniform input dimensions, the maximum length of enhancer and promoter sequences is set respectively. The insufficient part is padded with zero vectors, and the excessive part is truncated. The final output is a standardized model input tensor containing X_enhancer, X_promoter and label.

[0065] Furthermore, the model used in step S4 is based on the Transformer infrastructure, introducing a convolutional feature extraction module, a sine-cosine position encoding module, a cross-attention encoder layer, and a multi-layer feedforward network module to capture potential long-range regulatory relationships in the sequence. This model takes the transcription factor sequence (including binding order and binding strength) as input, learns the regulatory interaction representation between enhancers and promoters, and outputs the predicted probability value of the interaction. During the model training process, the binary cross-entropy loss function (Binary Cross-Entropy Loss) and the Adam optimizer were used for optimization. Combined with a ten-fold cross-validation strategy and multiple performance evaluation indicators (such as AUROC, AUPRC, etc.), the stability and robustness of the prediction effect were ensured. Through the visualization and explanatory modules of the attention weights, this model not only has good predictive capabilities, but also provides theoretical support for the subsequent revelation of the mechanism of combined action of regulatory factors. The specific steps are as follows:

[0066] (1) Input tensor

[0067] This model receives two input sequences , .in represents the enhancer sequence tensor, represents the promoter sequence tensor; and represent the padded lengths of enhancer and promoter sequences, respectively, and d represents the characteristic dimension of each site, including transcription factor one-hot coding and binding strength (d=27).

[0068] (2) Local feature extraction

[0069] Use two convolutional blocks to and Perform local feature extraction:

[0070] ;

[0071] ;

[0072] The same goes for Processed The convolution kernel sizes are 5 and 3 respectively, and the activation function is ReLU, followed by batch normalization (BN) and maximum pooling (MaxPooling) to improve the local sequence combination modeling capability.

[0073] (3) Positional Encoding

[0074] Since the convolution and attention mechanisms themselves do not have absolute position information, in order to enhance the sequential modeling capability, sine and cosine position encoding is introduced:

[0075] ;

[0076] ;

[0077] in is the position vector, It is the dimension, is the model dimension. This positional encoding vector is added to the convolution output:

[0078] ;

[0079] ;

[0080] (4) Cross-Attention

[0081] The cross-attention module is used to capture the potential regulatory dependencies between enhancer and promoter regions. A cross-attention operation is performed between the coding enhancer and promoter. Taking the enhancer-promoter focus as an example, let:

[0082] ;

[0083] ;

[0084] ;

[0085] Then the output of cross-attention is:

[0086] ;

[0087] This mechanism allows enhancer regions to dynamically focus on locations within the promoter region with the greatest regulatory potential.

[0088] The bidirectional attention operation is as follows:

[0089] ;

[0090] ;

[0091] Layer Normalization and Position-WiseFeedforward Network are added after each cross-attention output to enhance feature expression capabilities.

[0092] ;

[0093] The input vector X first undergoes nonlinear mapping through a feedforward network, maintaining the output dimension consistent with the input. This output is then residually connected to the original input, and layer normalization is performed again to effectively mitigate issues such as vanishing gradients and network degradation. Furthermore, to enhance the model's ability to model high-level semantic relationships, three layers of cross-attention encoder modules are stacked on each channel to further improve information fusion and abstraction capabilities.

[0094] (5) Output and classification

[0095] In the final stage of the model architecture, the feature maps of enhancers and promoters, after undergoing the cross-attention and encoder modules, are fed into a global average pooling layer, compressing sequence features of varying lengths into fixed-dimensional vector representations. Subsequently, the representation vectors of these two regions are concatenated into a single joint feature vector. A fully connected neural network module with dropout performs nonlinear transformation and feature fusion to enhance the model's discriminative power and prevent overfitting. Finally, the fused features pass through a sigmoid-activated output layer, generating a probability value between 0 and 1 indicating whether the enhancer-promoter pair has a true regulatory effect. When this value exceeds a preset threshold, the model predicts the presence of a potential long-range regulatory relationship.

[0096] In-depth, step S5 performs model training and evaluation. During the training phase, the Stratified K-Fold Cross-Validation method is used to divide the data. In each round of iteration, the ratio of positive and negative samples in the training set and the validation set is kept consistent, thereby effectively alleviating the bias caused by class imbalance at the data level. In terms of loss function, weighted binary cross-entropy is used to automatically assign class weights based on the ratio of positive and negative samples in each round of training, further enhancing the stability of the model when processing class-imbalanced data. In terms of optimizer, the Adam optimizer is selected, and the initial learning rate is set to 1e-3. The model parameters are dynamically updated during the training process. In order to comprehensively evaluate the performance of the model, a multi-dimensional indicator system is set, including accuracy, area under the receiver operating characteristic curve (AUROC), and area under the precision-recall curve (AUPRC).

[0097] To evaluate the performance and versatility of the model, this solution was experimentally validated on two human cell lines, K562 and GM12878, which are widely used in gene regulation research. First, a ten-fold cross-validation experiment was conducted on the K562 and GM12878 datasets to assess the stability and robustness of the model on different subsets. Figure 3 The results of a 10-fold cross-validation analysis of the K562 cell line are shown. The left figure shows the receiver operating characteristic (ROC) curve for each fold, and the right figure shows the predictive response (PR) curve. The model performed reliably across all folds, with an average AUROC of 0.927 and an AUPRC of 0.939. The relatively stable performance across folds demonstrates the model's strong ability to identify true enhancer-promoter pairings. Figure 4 The results of ten-fold cross-validation for the GM12878 cell line are shown. The left figure shows the receiver operating characteristic (ROC) curve for each fold, and the right figure shows the predictive response (PR) curve. In the GM12878 cell line, the proposed model maintains high generalization and recognition accuracy, with an average AUROC of 0.924 and an AUPRC of 0.934, demonstrating that the proposed method is equally applicable across different cell types.

[0098] In order to further verify the advantages of this method, the classic convolutional neural network (CNN) structure and the TargetFinder project method were selected as comparison models, and comparative experiments were carried out under balanced sampling and original unbalanced data settings.

[0099] Figure 5 The comparison results of downsampled data from the K562 cell line are shown, with the ROC curve on the left and the PR curve on the right. The model achieved an AUROC of 0.927 and an AUPRC of 0.939, significantly outperforming CNN (AUROC=0.906, AUPRC=0.913) and TargetFinder (AUROC=0.911, AUPRC=0.903).

[0100] In the case of the original imbalanced data, due to the large amount of data, 5-fold cross validation was selected. This method still showed good minority class recognition ability (AUPRC = 0.706), while CNN and TargetFinder were 0.543 and 0.444 respectively. Figure 6 shown.

[0101] In the GM12878 cell line, the AUROC of this method on the downsampled dataset is 0.924 and AUPRC is 0.934, which is also better than CNN (AUROC=0.901, AUPRC=0.905) and TargetFinder (AUROC=0.902, AUPRC=0.894). Figure 7 shown.

[0102] In the original imbalanced samples, the Transformer model still achieved good performance (AUPRC=0.650), while CNN and TargetFinder were 0.623 and 0.510 respectively. Figure 8 shown.

[0103] Finally, step S6 uses the cross-attention mechanism within the model to conduct in-depth interpretability analysis of the transcription factor interactions between enhancer and promoter regions, revealing the biological significance of the regulatory mechanism.

[0104] Figure 9 This is the K562 cell line enhancer-promoter sample attention heat map.

[0105] Figure 10 This is the attention heat map of enhancer-promoter samples in the GM12878 cell line.

[0106] Figure 11 The high-frequency interactions of transcription factor combinations with attention values ​​greater than 0.5 in positive samples are shown, and the frequencies of the top 50 combinations are visualized. The left figure is for the K562 cell line and the right figure is for the GM12878 cell line.

[0107] Figure 12 The frequency distribution of TF in enhancer and promoter regions in the high-attention region of the positive sample of K562 cell line is shown, where the left figure is the enhancer region and the right figure is the promoter region.

[0108] Figure 13 The frequency distribution of TF in enhancer and promoter regions in the high-attention region of the GM12878 cell line positive sample is shown, with the left figure showing the enhancer region and the right figure showing the promoter region.

[0109] Figure 14 、 Figure 15 The directed regulatory networks of transcription factors constructed based on the model attention weights of positive samples of K562 cell lines and GM12878 cell lines are respectively displayed, and the potential regulatory module structure is revealed through community division.

[0110] In summary, this application combines the spatial order and binding strength information extracted from ChIP-seq peaks to construct a transcription factor sequence feature representation, and introduces the Cross-Attention mechanism in the Transformer architecture to effectively model the long-range regulatory relationship between enhancers and promoters, thereby achieving high-precision prediction of enhancer-promoter interactions.

[0111] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Persons skilled in the art may make various modifications, additions, or substitutions to the described specific embodiments without departing from the spirit of the present invention or exceeding the scope of the appended claims.

[0112] Although the term "cross attention" is used more frequently in this article, the possibility of using other terms is not excluded. These terms are used only to more conveniently describe and explain the essence of the present invention; interpreting them as any additional limitations is contrary to the spirit of the present invention.

Claims

1. A prediction method for enhancer-promoter interactions based on transcription factor cross-attention, which uses a transformer model based on transcription factor cross-attention to predict enhancer-promoter interactions, wherein the transformer model includes an input layer, the input layer is connected to a sine-cosine position encoding module via a convolutional feature extraction module, and the sine-cosine position encoding module is connected to a multi-layer feedforward network module via a cross-attention encoder layer, characterized in that: The steps include: S1: Data acquisition and preprocessing, building a data foundation for enhancer-promoter interaction prediction; S11: Download ChIP-seq data of 26 key transcription factors in GM12878 and K562 cell lines; S12: Obtain enhancer and promoter pairing information files, transcription start site information files, and GM12878 and K562 cell type training set files; S2: Standardization of enhancer and promoter regions, unified region definition and coordinate processing; S21: Enhancer processing, extending 3 kb to both sides of the enhancer center to generate a 6 kb normalized region; S22: promoter processing, integrating the transcription start site information and defining the transcription factor site ± 2.5 kb as the promoter region; S23: Data screening: retain valid samples with strand information matching within the enhancer-promoter pairing region, and remove samples with multi-to-multi matching and no strand matching; S3: Key transcription factor feature encoding and input tensor generation, converting key transcription factor binding information into structured input; S31: Parse the ChIP-seq file of transcription factors to extract the binding peak position and binding intensity of transcription factors; S32: Assign a unique character code to each transcription factor, build a unified coding dictionary and store it as a JSON file; S33: For each enhancer or promoter region, extract all overlapping transcription factor binding sites within its coordinate range, generate transcription factor coding sequences in positional order, and record the corresponding binding strength numerical values. If it is a negative strand promoter, its sequence is reversed; S34: One-hot encode the transcription factor sequence and add an additional column to represent the binding strength of the transcription factor at that position to generate a tensor representation; S35: Set the maximum length for enhancer and promoter sequences respectively, fill the insufficient part with zero vector, truncate the excessive part, and output the standardized model input tensor; S4: Model architecture design, building a transformer model that integrates local features and global interactions; S5: Model training and evaluation, optimization measurement and robustness verification; S6: We use the cross-attention mechanism within the transformer model to perform interpretable analysis of transcription factor interactions between enhancer and promoter regions.

2. A method for predicting enhancer-promoter interactions based on transcription factor cross-attention according to claim 1, characterized in that: The step S4 includes the following steps: S41: Input enhancer and promoter sequence tensors, unifying sequence length and site feature dimensions; S42: Extracting local sequential patterns using different convolution kernels, followed by ReLU activation, batch normalization, and max pooling; S43: Perform sine and cosine position encoding and feature fusion; S44: Perform bidirectional attention calculation, perform nonlinear mapping through the feedforward network, followed by residual connection and layer normalization.

3. The method for predicting enhancer-promoter interactions based on transcription factor cross-attention according to claim 2, characterized in that: The step S4 inputs the feature maps of the enhancer and promoter after passing through the cross-attention encoder layer into the global average pooling layer, compresses the sequence features of different lengths into vector representations of fixed dimensions, splices the representation vectors of the two regions into an overall joint feature vector, and performs nonlinear transformation and feature fusion through a fully connected neural network module with a Dropout mechanism. The fused features are passed through a Sigmoid-activated output layer to generate a probability value between 0 and 1. When the value is greater than a preset threshold, the model predicts that the enhancer-promoter has a potential long-range regulatory relationship.

4. The method for predicting enhancer-promoter interactions based on transcription factor cross-attention according to claim 2, characterized in that: The step S4 stacks three layers of cross-attention encoder layers on each channel.

5. The method for predicting enhancer-promoter interactions based on transcription factor cross-attention according to claim 1, characterized in that: In step S5, the data is divided by a ten-fold cross validation method during the training phase. The loss function uses weighted binary cross entropy combined with category weights to process unbalanced data. The Adam optimizer is used as the optimizer, and the initial learning rate is set to 1e-3 and the parameters are adjusted dynamically.

6. A method for predicting enhancer-promoter interactions based on transcription factor cross-attention according to claim 5, characterized in that: The step S5 sets a multi-dimensional indicator system to evaluate the model performance, including accuracy, area under the receiver operating characteristic curve, and area under the precision-recall curve.

Citation Information

Patent Citations

  • Enhancer-promoter interaction prediction method based on DNA sequence and genome signal characteristics

    CN117542413A

  • Laser radar odometer method based on inter-frame overlapping region point cloud

    CN117824699A