Method for analyzing and predicting activity of long-sequence enhancer
By combining Performer's DNA language model and advanced feature extraction tower, the shortcomings of the long-sequence enhancer prediction model in capturing long-distance dependencies and interpretability are solved, and efficient and accurate enhancer activity analysis is achieved, which improves the model's prediction ability and biological interpretability.
Patent Information
- Application Number
- CN202510603950.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-09-02
AI Technical Summary
When processing long-sequence data, existing enhancer prediction models are difficult to effectively capture long-range dependencies of complex motifs, and lack interpretability analysis of enhancer activity, resulting in poor performance in high-dimensional data and hierarchical tasks.
A long-sequence enhancer activity analysis prediction method is adopted, combining Performer-based DNA language model and advanced feature extraction tower, and through embedding modules, attention matrix improvement and high-resolution feature extraction, long-term dependencies are captured to achieve efficient prediction.
It significantly improves the prediction accuracy and interpretability of long-sequence enhancer activity, and can explore the diversity and complexity of enhancer research tasks at different scales, providing more comprehensive theoretical support and practical guidance.
Smart Images

Figure CN120581068A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of enhancer analysis, and in particular relates to a method for analyzing and predicting the activity of long-sequence enhancers. Background Art
[0002] Enhancers are DNA sequences within the genome that play a key role in gene regulation. Key to their function lies a key motif within them, determining their activity and the regulation of target genes. Enhancers can regulate gene expression not only by modulating transcription rates but also through mutations in enhancer sequences. For example, mutations that inhibit enhancer activity can reduce gene expression, while those that do not can increase it, potentially leading to abnormal development and disease. For example, inherited mutations that affect enhancer activity have been linked to intestinal inflammation, polydactyly, and even certain cancers. Some enhancers can even exert long-range influences on gene expression. Therefore, exploring the regulation of enhancer activity near and beyond target genes is crucial.
[0003] Existing enhancer prediction models have made some progress in interpretability. As research deepens, researchers are placing higher demands on the interpretability of model behavior, hoping to understand the underlying biological mechanisms. For example, they are exploring the biological basis for predictive models reaching specific thresholds and the mechanisms by which complex weight parameters in deep neural networks influence decision-making, despite the fact that these parameters often lack intuitive biological meaning. Furthermore, leveraging the knowledge base built from the model's prior learning can predict potential unknown biological mechanisms or generate innovative hypotheses. Through systematic a posteriori analysis, they provide scientists with efficient tools to facilitate in-depth exploration in bioinformatics. In the field of interpretable machine learning (IML), various methods have been applied to the interpretation of enhancer prediction models. For example, interpretability research based on traditional machine learning models such as XGBoost and random forests extracts additional feature information to explain the model's prediction scores or classification results, thereby enhancing understanding of model behavior. However, IML methods have limitations, such as only capturing partial feature interaction information and relying on externally added features rather than intrinsic properties of the model itself. For high-dimensional data or hierarchical tasks, interpretability methods based on deep learning generally perform better. For example, the attention mechanism can be used to obtain the weight distribution of different sequence positions, thereby identifying key regions, especially for enhancer sequences containing important motifs.
[0004] The deep learning-based model SMFM uses an attention mechanism to perform feature modeling and motif analysis on enhancers within a 200-bp window, but is limited to short-range regions. DeepCAPE, on the other hand, predicts enhancers using 300-bp sequences and reveals motif importance through visualization convolutional layers. The EPCOT framework uses weights to represent the influence of motifs on enhancer activity, emphasizing their importance. However, it fails to quantify the actual contribution of base positions, limiting its ability to evaluate and interpret undiscovered motifs and large-scale sequence tasks. This problem remains unresolved. Currently, popular model interpretability methods focus on identifying and analyzing short enhancer sequences to discover relevant motifs. When dealing with data spanning hundreds or thousands of bases, the analysis and interpretation of complex motifs still faces both broad prospects and significant challenges. This field urgently requires more advanced interpretability methods that can address the complexity and diversity of motifs in long sequence data, thereby providing more comprehensive theoretical support and practical guidance for the functional elucidation of enhancers and other genomic elements. Summary of the Invention
[0005] In response to the above-mentioned deficiencies in the prior art, the present invention provides a method for analyzing and predicting enhancer activity in long sequences, which solves the problem that current models cannot well explore the sensitivity of sequence scale to task performance.
[0006] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is: a method for analyzing and predicting the activity of enhancers in long sequences, comprising the following steps:
[0007] S1, obtain the DNA sequence, embed the input sequence into the module, and generate the embedded data;
[0008] S2. Input the embedded data into the DNA language model to generate enhancer sequence information features;
[0009] S3, input the enhancer sequence information features into the advanced feature extraction tower to generate high-resolution features;
[0010] S4. Input the high-resolution features into the prediction module to generate high-resolution enhancer activity prediction results.
[0011] Further: S1 is specifically:
[0012] Enhancer activity data based on STARR-seq of five cell lines were obtained from the ENCODE database, and DNA sequences were extracted from the enhancer activity data using GKMSVM.
[0013] Furthermore: In S1, the method for generating embedded data is specifically:
[0014] A1. Input the DNA sequence into the Embedding layer to obtain the embedded output of the DNA sequence;
[0015] A2. Add position encoding to the embedded output of the DNA sequence to generate embedded data.
[0016] Furthermore, in A1, the expression of the embedded output Y of the DNA sequence is specifically:
[0017] Y=[e1,e2,...,e n ,...]
[0018] Where, e n is the embedding vector, n is the index of each dimension in the embedding vector, n∈[0,d-1], e n =E[x n ,:],x n is the index of each base position in the DNA sequence, E is the corresponding embedding matrix, E∈R V×d , V is the embedding table size, d is the dimension of the embedding vector;
[0019] In A2, the expression for generating the embedded data X is:
[0020] X=Y+PE
[0021] Where PE is the position encoding matrix;
[0022]
[0023] Where pos is the position in the sequence, even is an odd number, and odd is an even number.
[0024] Furthermore: In S2, the method for generating enhancer sequence information features is specifically:
[0025] The improved attention matrix is calculated based on the embedded data through the DNA language model to extract the information features of the enhancer sequence. The expression of the improved attention matrix Att is specifically:
[0026] Att=D -1 (Q′(K′) T V)
[0027] D=diag(Q′(K′) T 1 l )
[0028] Where Q is the first linear transformation parameter, Q = W Q X, W Q is the first weight matrix parameter, K is the second linear transformation parameter, K = W K X, W K is the second weight matrix parameter, V is the third linear transformation parameter, V = W VX, W V is the third weight matrix parameter, 1 l is the length of the uniform vector, T is the transpose sign, D is the intermediate matrix for calculating attention, diag(·) is the diagonal matrix with the input vector as the diagonal, Q′ is the random feature of Q, K′ is the random feature of K, is the characteristic function.
[0029] Furthermore: in S3, the advanced feature extraction tower includes several Stem blocks connected in sequence, and each Stem block includes a convolutional layer, a residual connection layer, a first batch normalization layer and a maximum pooling layer connected in sequence;
[0030] The convolutional layer is used to gradually extract the features of the input information, the residual connection layer is used to alleviate the gradient disappearance, and the batch normalization layer is used to alleviate the gradient disappearance and gradient explosion phenomena.
[0031] Further: In S4, the prediction module includes several component modules, each component module includes an input layer, a second batch normalization layer and an output layer connected in sequence.
[0032] Furthermore, in S4, the expression of the average loss L of the prediction module is specifically:
[0033]
[0034] Where N is the total number of samples of different tasks, M is the length of the output signal, is the real signal, is the predicted signal, and MSE(·) is the error loss function between the predicted value and the true value.
[0035] The beneficial effects of the present invention are:
[0036] (1) This paper provides a method for analyzing and predicting enhancer activity over long sequences. It innovatively designs a framework that combines an advanced language model based on Performer with a high-resolution feature extractor. This framework introduces an improved Transformer architecture to efficiently embed learning features and extract long-range dependencies. Combined with a high-resolution feature extractor, this framework achieves efficient prediction capabilities, including enhancer activity prediction, sequence scaling, and enhancer prediction task sensitivity studies. Through interpretable analysis, critical motifs in enhancers were discovered from multiple perspectives, demonstrating their effectiveness.
[0037] (2) The framework proposed in this paper cleverly handles the ability to capture the characteristics of long sequence enhancer sequence data, and at the same time explores the different expressiveness of sequences on enhancer research tasks at multiple scales, providing better performance than current methods.
[0038] (3) Based on the comparative experiments, ablation experiments, and interpretability analysis designed in the technical solution, the framework proposed in this invention achieved leading results. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 This is a flow chart of a method for analyzing and predicting long sequence enhancer activity of the present invention.
[0040] Figure 2 A performance comparison chart to explore the impact of sequence data on tasks under scene data of different scales.
[0041] Figure 3 Performance comparison chart to explore the best architecture of the model.
[0042] Figure 4 Validation plot across cell types.
[0043] Figure 5 A graph of sequence motifs extracted from five cell types for the DNA language model.
[0044] Figure 6 Schematic diagram of high-contribution sequence motifs for LEAP to reveal cell type-specific enhancers Figure 1 .
[0045] Figure 7 Schematic diagram of high-contribution sequence motifs for LEAP to reveal cell type-specific enhancers Figure 2 . DETAILED DESCRIPTION
[0046] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.
[0047] like Figure 1 As shown, in one embodiment of the present invention, a method for analyzing and predicting long sequence enhancer activity includes the following steps:
[0048] S1, obtain the DNA sequence, embed the input sequence into the module, and generate the embedded data;
[0049] S2. Input the embedded data into the DNA language model to generate enhancer sequence information features;
[0050] S3, input the enhancer sequence information features into the advanced feature extraction tower to generate high-resolution features;
[0051] S4. Input the high-resolution features into the prediction module to generate high-resolution enhancer activity prediction results.
[0052] In this embodiment, the present invention proposes a Performer-based long sequence enhancer activity prediction and analysis framework, named LEAP, which aims to capture more distant enhancer information features to achieve efficient and accurate prediction of enhancer activity, as well as explore richer internal mechanism principles. The framework mainly consists of four parts: sequence embedding module (SEM), DNA language model (DNA-LM), advanced feature extraction tower (Stem Tower), and prediction module (FPM). The sequence embedding module efficiently captures long DNA sequence data, effectively embeds high-scale data features, and retains the original information. The DNA language model is based on the advanced paradigm of Performer and customizes a proprietary architecture to effectively solve the attention mechanism dilemma of long sequence scenarios of enhancers. The advanced feature extraction tower is to extract more efficient features, further extract and reduce the features, and build a flexible model structure backbone tower. The prediction module constructs multiple neuron layers to achieve high-resolution enhancer activity prediction and obtain the final result.
[0053] S1 is specifically:
[0054] Enhancer activity data based on STARR-seq of five cell lines were obtained from the ENCODE database, and DNA sequences were extracted from the enhancer activity data using GKMSVM.
[0055] In this example, for enhancer activity data from STARR-seq, negative sequences with similar GC content to the positive sequences were selected. The peaks were then extended by 500 bp on either side to generate a 1:1 balanced dataset of 1001 bp positive and 1001 bp negative sequences. For activity signals based on positive and negative samples, deepTools was used to extract the positive and negative sample signals from the bigwig files. The data was normalized using a log2(1 + signal) transformation to reduce bias and used for training and validation of LEAP.
[0056] This paper employs a dual embedding method, using an embedding layer and positional encoding, to embed input data as a 200-dimensional vector, facilitating the subsequent performer to extract high-level gene sequence features. This not only effectively embeds enhancer-specific encodings but also captures knowledge about sequences and even longer distances, while also acquiring global positional information about sequences. Leveraging prior knowledge, this approach reduces the difficulty of model training. The embedding process is described below:
[0057] A1. Input the DNA sequence into the Embedding layer to obtain the embedded output of the DNA sequence;
[0058] A2. Add position encoding to the embedded output of the DNA sequence to generate embedded data.
[0059] In this embodiment, compared to the limitations of the receptive field of other models, LEAP exponentially expands the receptive field of features, obtains more feature information for embedding, and efficiently achieves the target task. For enhanced subsequence data, in order to convert a single continuous variable into an embeddable vector, an effective way of mapping phrases, sentences, expressions, etc. in NLP to vector space is adopted - Embedding. Compared with the limited token embedding and sparse representation of One-hot vectors in the language model BERT, Embedding with high semantic expression ability can map the tags of variables to a continuous, dense, low-dimensional vector space, so that the model can better understand the semantic relationship between variable tags, retain position information communication between longer distances, and retain more feature information while densely embedding vectors, which enables the model to capture the required features more effectively.
[0060] In A1, the expression of the embedded output Y of the DNA sequence is specifically:
[0061] Y=[e1,e2,...,e n ,...]
[0062] Where, e n is the embedding vector, n is the index of each dimension in the embedding vector (A>0, C>1, G>2, T>3), n∈[0,d-1], e n =E[x n ,:],x n is the index of each base position in the DNA sequence, E is the corresponding embedding matrix, E∈R V×d , V is the embedding table size, d is the dimension of the embedding vector;
[0063] In A2, the expression for generating the embedded data X is:
[0064] X=Y+PE
[0065] Where PE is the position encoding matrix;
[0066]
[0067] Where pos is the position in the sequence, even is an odd number, and odd is an even number.
[0068] In this example, positional encoding, leveraging the periodic characteristics of sine and cosine functions, can capture sequence information at different scales, enabling the model to flexibly learn semantic relationships based on positional variations. Consequently, the model can better understand the internal meaning of the enhancer sequence.
[0069] In S2, the method for generating enhancer sequence information features is specifically as follows:
[0070] The DNA language model is used to calculate the improved attention matrix based on the embedded data to extract the information features of the enhancer sequence;
[0071] In this embodiment, in order to obtain a wider global receptive field, the present invention constructs a DNA language model based on Performer. The core improvement of Performer is to use matrix decomposition combined with an approximate method based on random feature mapping to reduce the attention calculation complexity from O(n 2 ) is reduced to Q(n), replacing the traditional dot product calculation mechanism. This improvement enables the model to process long sequence data more efficiently, while expanding the scale of the receptive field to capture more enhanced subsequence information. The expression of the improved attention matrix Att is specifically:
[0072] Att=D -1 (Q′(K′) T V)
[0073] D=diag(Q′(K′) T 1 l )
[0074] Where Q is the first linear transformation parameter, Q = W Q X, W Q is the first weight matrix parameter, K is the second linear transformation parameter, K = W K X, W K is the second weight matrix parameter, V is the third linear transformation parameter, V = W V X, W V is the third weight matrix parameter, 1 l is the length of the uniform vector, T is the transpose sign, D is the intermediate matrix for calculating attention, diag(·) is the diagonal matrix with the input vector as the diagonal, Q′ is the random feature of Q, K′ is the random feature of K, is the characteristic function, which is defined as:
[0075]
[0076] Where c is a positive constant, m is the dimension of the matrix, ω is the random feature matrix, x is the input of the feature function, and f(·) is the matrix mapping function.
[0077] In S3, the advanced feature extraction tower consists of several Stem blocks connected in sequence. Each Stem block consists of a convolutional layer, a residual connection layer, a first batch normalization layer, and a maximum pooling layer connected in sequence.
[0078] The convolutional layer is used to gradually extract the features of the input information, the residual connection layer is used to alleviate the gradient disappearance, and the batch normalization layer is used to alleviate the gradient disappearance and gradient explosion phenomena.
[0079] In this embodiment, in order to facilitate the calculation of subsequent convolutional layers and reduce complexity, each Stem block includes a convolutional layer, a residual connection layer, a batch normalization layer and a maximum pooling layer. Multiple convolutional layers are used to gradually extract low-level to high-level features. The hierarchical design of different convolution kernel sizes balances the model to obtain local information to global information, and gradually enhances the feature expression capability. For complex neural network training, residual connections can alleviate gradient vanishing, retain a better input information flow, and speed up the convergence time of the model to make the task more efficient. The batch normalization layer effectively alleviates the gradient vanishing and gradient explosion phenomena, thereby accelerating the model convergence rate, and introducing an appropriate amount of noise to improve the model generalization ability, thereby alleviating the problem of internal covariate offset. In addition, by integrating the maximum pooling operation, the computational cost is further reduced.
[0080] In S4, the prediction module includes several component modules, each of which includes an input layer, a second batch normalization layer, and an output layer connected in sequence.
[0081] In this example, the output of the advanced feature extraction tower first passes through the Flatten layer before being used in the prediction module to obtain the final result. The output channel is 1001, representing the activity scale of the base-resolution enhancer. The prediction module is internally configured with multiple neuron layers, with each component module consisting of an input layer, a batch normalization layer, and an output layer.
[0082] In S4, the expression of the average loss L of the prediction module is specifically:
[0083]
[0084] Where N is the total number of samples of different tasks, M is the length of the output signal, is the real signal, is the predicted signal, and MSE(·) is the error loss function between the predicted value and the true value.
[0085] In this embodiment, the present invention also provides comparative experiments, ablation experiments, transcellular experiments, and multi-angle analysis interpretability experiments, as shown below:
[0086] Comparative experiment
[0087] The performance of our proposed framework for predicting enhancer activity was compared with seven published methods. We applied the same five cell-specific enhancer datasets to ensure a fair comparison. Table 1 shows the comparison of LEAP with all baselines for all two metrics across the five cell-based datasets.
[0088] Table 1 Comparison with other methods on simulated datasets
[0089]
[0090]
[0091] LEAP demonstrated superior performance in predicting enhancer activity, outperforming all compared methods. This study demonstrated that the LEAP framework significantly improves enhancer activity prediction accuracy by integrating efficient embedding algorithms, cutting-edge language modeling techniques, and a high-resolution feature extraction architecture. Compared to all baselines, LEAP achieved average improvements in PCC (0.044, 0.059, 0.073, 0.043, and 0.049) and average reductions in MSE (0.127, 0.138, 0.148, 0.117, and 0.114) across the five datasets.
[0092] Compared to other cell datasets, LEAP performed the worst on the K562 cell dataset, with PCC and MSE of 0.772 and 0.459, respectively. Most other baselines also exhibited this trend, which we attribute to the size of the dataset. However, the model performed best on the A549 cell dataset, which had the largest dataset. PCC and MSE for A549 were 0.803 and 0.378, respectively.
[0093] LEAP not only performs best compared to all other baselines, but also far surpasses DeepSTARR and Basenji on the original short sequence scenario tasks. We speculate that this may be due to the sequence scale. Therefore, we selected DeepSTARR, which has the most stable task performance, to explore whether increasing the scale of different sequence settings will improve its prediction level.
[0094] Models that perform well at a single scale may struggle with more complex sequence problems, such as those related to ordinary short sequence enhancers and longer sequence super enhancers. By analyzing the performance of the model at different scale settings, we can quantitatively compare the relative importance of short-range local features and long-range local features. The performance on problems with different sequence lengths can also demonstrate the effect of model generalization. Obtaining more contextual information is also very helpful in revealing the functional dependence of enhancers. Compared with long distances, short sequences are more likely to lose some important contextual information, which causes certain difficulties for the analysis task. The sequence motifs in enhancers are unique functional characteristics of enhancers. This part of the region may exist in different positions. Therefore, the existence of specific motifs will become uncertain in different length ranges. In order to explore whether they may be distributed in a wider range of contextual information. The present invention has conducted research on feature acquisition in different scale ranges from short to long, and a deeper understanding of enhancer activity may be more comprehensive.
[0095] In order to explore the influence of the length of sequence data on the performance of enhancer activity prediction tasks based on LEAP, this paper sets up experimental analysis at different scales and selects DeepSTARR, which has an advantage in short-scale tasks, as a benchmark. In five types of cell data, each DNA sequence is divided into 125bp data enhancement sliding windows, and activity prediction experiments are carried out at different scales of 249bp, 375bp, 500bp, 625bp, 750bp, 875bp, and 1001bp. Figure 2 Figure (a) shows a control experiment using the same data with different scale settings as input, starting from the center of the original sequence length and divided into lengths of 249 bp to 1001 bp. Figure (b) shows a control experiment with DeepSTARR on five cell datasets, each of which was divided into seven different scale settings. Blue represents LEAP and red represents DeepSTARR. Values are Pearson correlation coefficients.
[0096] Across all cell datasets of enhancer sequences divided into seven scales, LEAP achieved the best results in prediction tasks, particularly at longer distances. Using short, medium, and long sequence lengths as the evaluation criteria, LEAP achieved high average task performance (0.014 and 0.023) for relatively short sequences of 249bp and 375bp; high average task performance (0.023 and 0.043) for medium-length sequences of 500bp and 625bp; and high average task performance (0.033, 0.044, and 0.047) for longer sequences of 750bp, 875bp, and 1001bp. Experiments demonstrate that LEAP not only outperforms DeepSTARR on short sequences of 249bp and 375bp, but also performs even better on medium-length and longer sequences of 500bp to 1001bp. LEAP is effectively applicable to problems at different scales. The experiment also proved that compared with the initial distance setting, the performance of the two tasks will improve with the increase of scale, and the overall trend is upward. Capturing the features of enhancers at a longer distance has a significant improvement on the task effect. There may be information exchange between distant enhancers that is helpful for analysis. However, DeepSTARR's task performance has reached a bottleneck when it reaches about 625bp, while LEAP can reach a longer distance of 1001bp, obtaining more feature acquisition opportunities and the advantage of information dependence. The experimental analysis of multiple sequence lengths provides a certain strategic analysis for subsequent related long-distance enhancer interaction issues, aiming to be helpful for subsequent data fusion at different scales. The practice of this invention provides a reference for future related theoretical research.
[0097] Ablation experiments
[0098] To deeply analyze the structural optimization potential of the LEAP framework, we systematically conducted multiple rounds of ablation experiments. The experimental design followed the principle of single-variable adjustment to effectively control computing resource consumption, and non-target variables were uniformly set to default parameters in the experiments.
[0099] (1) The current model can capture the long-range dependencies of short-scale enhancer sequences, but its performance on longer-range enhancer tasks is still insufficient, and it cannot fully explore the sensitivity of sequence scale to task performance. One-hot encoding (LEAP-ONE) is selected as the embedding method: the original embedding method of LEAP is modified to use the commonly used one-hot encoding method for sequence embedding.
[0100] (2) Not using the DNA-LM module composition architecture (LEAP-LM): Remove the DNA-LM module and directly pass the embedded results to the subsequent modules.
[0101] (3) No Stem Tower module composition architecture (LEAP-ST): Remove the Stem Tower module and directly pass the output of the DNA-LM module into the subsequent modules.
[0102] (4) Complete LEAP framework composition (LEAP-ALL): Enhancer activity prediction is performed using the LEAP framework with a complete architecture.
[0103] The performance comparison is done with several different setups, such as Figure 3 Figure (a) shows the performance comparison of different settings on five datasets, using the Pearson correlation coefficient as the metric; (b) shows the performance comparison using the mean squared error as the evaluation metric. First, the effects of the embedding layer and one-hot encoding on LEAP were compared. One-hot encoding has been a recently effective embedding method, sharing the same properties as embedding. While it performs well in small-scale, simple tasks, its performance is limited when dealing with complex DNA sequence relationships, and its computational efficiency may be affected by the high-dimensional sparsity of the embedding vector. The embedding layer represents each base as a low-dimensional, dense vector by learning a smaller vector space, reducing memory usage and improving computational efficiency. Furthermore, inter-base relationships are crucial in biological contexts, but one-hot encoding does not consider them. Embedding, on the other hand, considers potential relationships or grammatical structures between bases, effectively capturing high-level features of DNA sequences. By mapping bases with similar functions or physical properties to similar vector spaces, it better captures relationships between bases. Furthermore, compared to the static one-hot encoding method, the Embedding layer continuously learns contextual semantic information during training to achieve self-adjustment. LEAP-ALL outperforms LEAP-ONE, with a higher PCC (0.039) and lower MSE (0.041).
[0104] Then, the present invention studies the impact of the DNA-LM module on performance. The Performer in the DNA-LM module has the advantage of a large receptive field, which can effectively utilize the global information in the sequence data and capture remote features to learn global representations. Using the Performer, the present invention can obtain a better information flow of enhancer elements of remote regulatory genes in the focus sequence while effectively retaining local information, and can effectively integrate their information to achieve model improvement. In addition, the attention mechanism in the Performer can effectively strengthen the focus on key features and optimize the attention allocation of local unimportant information. Enhancer sequence data is specifically encoded, and the sequence motifs composed of enhancer syntax are specific and independent in different cell types. Therefore, the allocation of importance weights for this unique region is also an effective means to solve the enhancer problem. In regional information with uneven feature distribution, the sequence relationship will be very complex, and the prediction accuracy will also be greatly tested. The flexible attention mechanism can have a great space to play in the unique enhancer activity information. Compared with other different architectural settings, the performance of the model is the lowest when DNA-LM is not available. LEAP-ALL improves PCC (0.052) and MSE (0.056) compared to LEAP-LM. DNA-LM is an effective choice for predicting enhancer activity.
[0105] In addition, the present invention compares the performance difference between using Stem Tower and not using it. For deep network models, convolution kernels are generally used to learn features, but some useful features may be lost during learning, and model training will become increasingly difficult to improve. The introduction of the residual connection module can improve the efficiency of gradient propagation along the residual connection line to effectively solve this problem. The composition of Stem Tower, combining multiple advantages, can not only capture relevant information through local receptive fields, but also has better extraction of details and global information, and can also transmit from low levels to high levels for feature fusion. The flexible module composition increases the ability to capture information of different lengths, which is beneficial to the model's multiple scale tasks. The model is also more stable for complex data scenarios and has more generalization robustness. LEAP-ALL is better than LEAP-ST, with PCC increased (0.039) and MSE reduced (0.046). Having a Stem Tower module is also a key component for accurately predicting enhancer activity.
[0106] Transcellular experiments
[0107] To investigate the generalization ability of the LEAP framework across different cell types, cross-cell type experiments are needed. A model trained on one of the five cell types is used to test the remaining four to examine whether the cell type-specific model is transferable to other cell type tasks. Figure 4 As shown, Figure 4The heatmap in the middle shows the performance of cross-cell type validation. The numbers in the rectangles represent the specific values of the performance metrics. The vertical axis represents the dataset, and the horizontal axis represents the source of the model parameter settings. The average PCC for cross-cell testing reached 0.687, which is 0.107 lower than that of the task-specific model. The average MSE was 0.538, which is 0.132 higher than that of the task-specific model. This shows that there is a certain degree of independence between different cell types. Specific enhancers in specific cells account for the majority, while shared enhancers may also exist between them. Experimental results show that LEAP has good generalization and still performs well in cross-cell tasks, but the best prediction performance is achieved when using specific models.
[0108] Analyzing interpretability experiments from multiple perspectives
[0109] To explore the correlation between enhancer sequence motifs and activity, we mined the internal syntax of enhancers. This study was conducted from two perspectives: (1) by obtaining attention weights, we mined and discovered motifs within specific cells, and (2) by using the perspective of mutant bases, we studied the contribution of important bases within enhancers to the prediction task. By finding that sequence motifs appear simultaneously at the near and far ends, we once again demonstrated the necessity of studying the long sequence scenario of the enhancer problem and proved that the LEAP analysis grammar principle has good interpretability.
[0110] In order to prove that the attention mechanism used has a great contribution to enhancer prediction, the present invention uses the attention mechanism existing in the language model to explore the internal semantics of enhancers. Since attention can pay special attention to patterns that have a key contribution to enhancer activity analysis, by extracting the larger attention weights in the encoder layer, the motifs that have made important contributions composed of enhancer syntax are found, thereby verifying whether the attention layer can learn motif information. By using the Performer-based LEAP as a template, the weight of the attention score is obtained to explain the content learned by the model. The present invention successfully discovered specific sequence motifs for different cell types. Figure 5Two important and significant sequence motifs captured in each of these data sets are shown. The present invention observes that for different cell types, the sequence motifs are also different from each other, that is, the syntax of enhancers is that there is specific independence, which is also meaningful for the identification of enhancers. The present invention compares the motifs found with the database in the widely used motif discovery tool TomTom, and finds that the matching motifs of A549 cells are (ELF3) and (TFAP4_DBD), the matching motifs of HCT116 cells are (ZNF768) and (RELB), the matching motifs of HepG2 cells are (Meis2_DBD_1) and (Zic1::Zic2), the matching motifs of K562 cells are (Spi1) and (Foxn1), and the matching motifs of MCF-7 cells are (HIC2) and (KLF4). These important patterns are of great significance for studying the internal principles and regulatory mechanisms that affect enhancer activity.
[0111] Although research from the perspective of attention can discover the key components of enhancers, little is known about the importance of enhancer bases. It is also very important to think about the importance contribution of enhancer bases, which has important research significance for the principle of information exchange within enhancers. In order to better understand the key factors affecting enhancer information expression, the present invention uses LEAP to analyze the impact of base mutations on enhancer activity, so that important sequence elements can be discovered. By using the latest analysis tools to conduct experiments, it was found that the high contribution expression score regions of nucleotides in the sequence can correspond to existing enhancer sequence motifs. The present invention uses TangerMEME to analyze enhancer sequences based on LEAP. After pre-training, LEAP predicts enhancer activity from the sequence. TangerMEME is an open source project for biological sequence analysis. It obtains the model's interest in a given DNA sequence by performing the silicon saturation mutation (ISM) function. The ISM function is used to replace each base in a given sequence with all other possible bases in sequence, and each base in the target region generates a corresponding attribution score, such as Figures 6-7 LEAP reveals highly contributing sequence motifs at cell-type-specific enhancers, displaying the attribution score for each position in the sequence. Figure 6 In the visualization, heatmap values for A, T, C, and G at different positions are displayed, with bases of high importance at each position highlighted (A in green, T in red, C in blue, and G in yellow). The height of the display refers to the contribution score, and the focus area corresponds to the heatmap display. A specific magnified image is shown at the bottom. The present invention discovered two highly prominent regions in the region (chr16:89571754-89572754), in which five similar known TF motifs were found. Figure 7In the analysis, two highly prominent regions were found in the (chr17:73323660-7324660) region, in which a total of seven similar known motifs were found.
[0112] Then, the difference in the magnitude of the change in the model prediction output before and after the replacement is evaluated. This difference in magnitude can be regarded as a measure of the importance or attribution of the model to its degree of interest. The larger the value, the greater the impact of the characteristic change after the mutation on the prediction, that is, the more important it is. Through ISM, we can discover and find the highly expressed nucleotide blocks on the target enhancer sequence. Below the heat map, the specific contribution scores are displayed for different bases at each position, and the bases with the most important influence are highlighted. Zoom in and observe the sequence position area with the most outstanding contribution, use a red rectangular box to mark the important and prominent area in the heat map and the area of the corresponding base independent high contribution score map, and enlarge and map them. The fragment sequence with a high prominent score is marked in the dotted box, and the ID and name label of the corresponding similar known motif in the TomTom database are indicated by arrows.
[0113] Specifically, enriched motif sequences in the database were found in the sequence data testing experiments for chr16:89571754-89572754 and chr17:73323660-7324660. The 170-190bp, 330-340bp, and 520-580bp regions of chr16 contained five motifs: MA0622.1 (Mlxip), MA0004.1 (Arnt), MA0863.1 (MTF1), MA0646.1 (GCM1), and UP00UP00036_2 (Myf6_secondary). The MA0004.1 (Arnt) motif appeared at different positions around 440bp and 770bp, respectively. Sequence fragments with high contribution scores were found at positions 150-250bp and 470-570bp of chr17. There are seven possible corresponding TF motifs, namely (MA1101.2 (BACH2), MA1117.1 (RELB), MA0006.1 (Ahr::Arnt), UP00104_1 (Hmx1_3423.1), MA1928.1 (BNC2), MA0106.3 (TP53), ZNF306_full).
[0114] In summary, this study explores the contributions of nucleotides at different positions and uses this to discover motifs. These motifs are located as far apart as 400 bp. Therefore, a wider receptive field, after acquiring more comprehensive global information, is irreplaceable in extracting sequence motifs. Although widely used, small-scale sequence analysis tasks may not perform well in downstream tasks such as interpretability. Long-scale sequence model analysis may address this issue. LEAP can effectively detect the presence of distal sequence motifs that other short-sequence models cannot detect, and can efficiently locate their specific locations with good fitness. ISM mutagenesis experiments also revealed that motifs within enhancers contribute significantly to the prediction of enhancer activity, and LEAP also has high sensitivity for this important region. Furthermore, it was found that the same motif can be present at different positions in the sequence. For example, the MA0004.1 (Arnt) motif in chr16:89571754-89572754 is located at positions 440 bp and 770 bp, respectively, with the two locations separated by more than 300 bp. If we only analyze the first or second half of the position using a local receptive field, we will lose the important attention to this motif, and its important research potential will be ignored. Compared to short sequence tasks, long-scale tasks are more capable of discovering this type of hidden information that is difficult to find in sequences.
[0115] In the description of the present invention, it should be understood that the terms "center", "thickness", "upper", "lower", "horizontal", "top", "bottom", "inner", "outer", "radial", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only and cannot be understood as indicating or implying the relative importance or the number of technical features implicitly specified. Therefore, the features defined by "first", "second", and "third" may explicitly or implicitly include one or more of such features.
Claims
1. A method for analyzing and predicting long sequence enhancer activity, characterized in that: The following steps are involved: S1, obtain the DNA sequence, embed the input sequence into the module, and generate the embedded data; S2. Input the embedded data into the DNA language model to generate enhancer sequence information features; S3, input the enhancer sequence information features into the advanced feature extraction tower to generate high-resolution features; S4. Input the high-resolution features into the prediction module to generate high-resolution enhancer activity prediction results.
2. The method for analyzing and predicting long sequence enhancer activity according to claim 1, wherein: S1 is specifically: Enhancer activity data based on STARR-seq of five cell lines were obtained from the ENCODE database, and DNA sequences were extracted from the enhancer activity data using GKMSVM.
3. The method for analyzing and predicting long sequence enhancer activity according to claim 1, wherein: In S1, the method for generating embedded data is as follows: A1. Input the DNA sequence into the Embedding layer to obtain the embedded output of the DNA sequence; A2. Add position encoding to the embedded output of the DNA sequence to generate embedded data.
4. The method for analyzing and predicting long sequence enhancer activity according to claim 3, wherein: In A1, the expression of the embedded output Y of the DNA sequence is specifically: Y=[e1,e2,...,e n ,...] Where, e n is the embedding vector, n is the index of each dimension in the embedding vector, n∈[0,d-1], e n =E[x n ,:],x n is the index of each base position in the DNA sequence, E is the corresponding embedding matrix, E∈R V×d , V is the embedding table size, d is the dimension of the embedding vector; In A2, the expression for generating the embedded data X is: X=Y+PE Where PE is the position encoding matrix; Where pos is the position in the sequence, even is an odd number, and odd is an even number.
5. The method for analyzing and predicting long sequence enhancer activity according to claim 1, wherein: In S2, the method for generating enhancer sequence information features is specifically as follows: The improved attention matrix is calculated based on the embedded data through the DNA language model to extract the information features of the enhancer sequence. The expression of the improved attention matrix Att is specifically: That=D -1 (Q′(K′) T V) D=diag(Q′(K′) T 1 l ) Where Q is the first linear transformation parameter, Q = W Q X, W Q is the first weight matrix parameter, K is the second linear transformation parameter, K = W K X, W K is the second weight matrix parameter, V is the third linear transformation parameter, V = W V X, W V is the third weight matrix parameter, 1 l is the length of the uniform vector, T is the transpose sign, D is the intermediate matrix for calculating attention, diag(·) is the diagonal matrix with the input vector as the diagonal, Q′ is the random feature of Q, K′ is the random feature of K, is the characteristic function.
6. The method for analyzing and predicting long sequence enhancer activity according to claim 1, wherein: In S3, the advanced feature extraction tower consists of several Stem blocks connected in sequence. Each Stem block consists of a convolutional layer, a residual connection layer, a first batch normalization layer, and a maximum pooling layer connected in sequence. The convolutional layer is used to gradually extract the features of the input information, the residual connection layer is used to alleviate the gradient disappearance, and the batch normalization layer is used to alleviate the gradient disappearance and gradient explosion phenomena.
7. The method for analyzing and predicting long sequence enhancer activity according to claim 1, wherein: In S4, the prediction module includes several component modules, each of which includes an input layer, a second batch normalization layer, and an output layer connected in sequence.
8. The method for analyzing and predicting long sequence enhancer activity according to claim 7, wherein: In S4, the expression of the average loss L of the prediction module is specifically: Where N is the total number of samples of different tasks, M is the length of the output signal, is the real signal, is the predicted signal, and MSE(·) is the error loss function between the predicted value and the true value.