Enhancer activity prediction method and system based on bidirectional long short-term memory network

By using bidirectional long and short-term memory networks and multiple sequence coding methods in enhancer activity prediction, combined with epigenetic modified data, the problem of insufficient accuracy and data universality of enhancer activity prediction in the prior art is solved, and higher prediction accuracy and robustness are achieved.

CN120146097APending Publication Date: 2025-06-13SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510226844.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art does not perform well in the prediction of enhancer activity, especially in quantitative prediction and data universality, and the model generalization performance is not high.

Method used

The enhancer activity prediction method based on bidirectional long and short-term memory network is adopted, and DNA sequence encoding is performed through reverse complementary k-mer, mismatched k-mer and word vector methods, and combined with epigenetic modification data, the two-way long and short-term memory network captures the two-way dependence relationship of sequence characteristics and performs feature fusion to predict enhancer activity.

Benefits of technology

It effectively improves the accuracy of enhancer activity prediction, can achieve more accurate prediction under limited data conditions, and has good robustness and generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146097A_ABST
    Figure CN120146097A_ABST
Patent Text Reader

Abstract

The invention provides an enhancer activity prediction method and system based on a bidirectional long-short-term memory network, coding is performed through a reverse complementary k-mer method, a mismatching k-mer method and a word vector method, different view angle features are constructed, and a bidirectional dependency relationship of sequence features coded through the word vector method is captured by using the bidirectional long-short-term memory network; splicing features obtained after reverse complementation k-mer and mismatching k-mer coding and splicing are fused with the feature representation, and then the activity of the enhancer is predicted. According to the scheme, experiments prove that the accuracy of enhancer activity prediction can be effectively improved, more accurate prediction can be achieved under the limited data condition, and good robustness is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of bioinformatics, and in particular relates to a method and system for predicting enhancer activity based on a bidirectional long short-term memory network. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] Gene regulation is a complex process that is key to the normal functioning of biological systems. As an important type of regulatory element, enhancers play an important role in regulating gene expression by interacting with promoters and other regulatory regions. However, unlike promoters, the function of enhancers does not depend on their position and orientation relative to the target gene, which increases the complexity of their identification and characterization. Therefore, it is very important to quantitatively predict the functional activity of enhancers.

[0004] Machine learning methods, especially deep learning models, show good prospects in improving the accuracy of enhancer activity prediction. However, there are still several problems that need to be solved in the current models for enhancer activity prediction: first, these methods are not ideal in quantitatively predicting enhancer activity due to relatively simple sequence encoding schemes and relatively simple model architectures; second, these methods have only been tested in human species or fruit fly species, which limits the universality of their data sets and the model generalization performance is not high. Summary of the invention

[0005] In order to overcome the deficiencies of the above-mentioned prior art, the present invention provides an enhancer activity prediction method and system based on a bidirectional long short-term memory network, which can effectively improve the accuracy of enhancer activity prediction and can achieve more accurate prediction under limited data conditions, and has good robustness.

[0006] In order to achieve the above object, the present invention adopts the following technical solution:

[0007] In a first aspect, the present invention provides a method for predicting enhancer activity based on a bidirectional long short-term memory network, comprising:

[0008] Obtain DNA sequence data and epigenetic modification data;

[0009] The DNA sequence data is encoded by reverse complementary k-mer, mismatched k-mer and word vector methods respectively;

[0010] Dimensionally splicing the sequence features encoded by the reverse complementary k-mer and the mismatched k-mer to obtain spliced ​​features, and extracting the spliced ​​features through a sequence extraction module;

[0011] Based on the context information of the sequence features encoded by the word vector method, a feature representation is obtained; the feature representation and the epigenetic modification data are respectively subjected to feature extraction through a Word2Vec module based on a bidirectional long short-term memory network and an epigenetic modification module;

[0012] The output results of the sequence extraction module, the Word2Vec module, and the epigenetic modification module are fused, and the activity of the enhancer is determined according to the fusion result.

[0013] In a second aspect, the present invention provides an enhancer activity prediction system based on a bidirectional long short-term memory network, including:

[0014] An acquisition module, which is configured to: acquire DNA sequence data and epigenetic modification data;

[0015] An encoding module, which is configured to: perform encoding processing on the DNA sequence data through reverse complementary k-mer, mismatched k-mer, and word vector methods respectively;

[0016] A first extraction module, which is configured to: splice the sequence features encoded by reverse complementary k-mer and mismatched k-mer in dimension to obtain a spliced feature, and perform feature extraction on the spliced feature through a sequence extraction module;

[0017] A second extraction module, which is configured to: based on the context information of the sequence features encoded by the word vector method, obtain a feature representation; perform feature extraction on the feature representation and the epigenetic modification data through a Word2Vec module based on a bidirectional long short-term memory network and an epigenetic modification module respectively;

[0018] A prediction module, which is configured to: fuse the output results of the sequence extraction module, the Word2Vec module, and the epigenetic modification module, and determine the activity of the enhancer according to the fusion result.

[0019] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in the first aspect is completed.

[0020] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first aspect is completed.

[0021] In a fifth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described in the first aspect is implemented.

[0022] The above one or more technical solutions have the following beneficial effects:

[0023] In the present invention, encoding is performed through reverse complementary k-mer, mismatched k-mer and word vector methods to construct features from different perspectives, and a bidirectional long short-term memory network is used to capture the bidirectional dependencies of the sequence features encoded by the word vector method; the concatenated features obtained after encoding and concatenating through reverse complementary k-mer and mismatched k-mer are fused with the feature representations, and then the activity of the enhancer is predicted. The solution of the present invention has been proven by experiments to be able to effectively improve the accuracy of enhancer activity prediction, and can achieve more accurate prediction under limited data conditions, and has good robustness.

[0024] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0026] Figure 1 It is a block diagram of the enhancer activity prediction method in the first embodiment of the present invention;

[0027] Figure 2 It is a network structure diagram of each module in the prediction model in the first embodiment of the present invention;

[0028] Figure 3 In it, A - F are respectively the performance evaluation effect diagrams of different module fusions in the first embodiment of the present invention and the cell line DEV;

[0029] Figure 4 In it, A - B are respectively the PCC and SCC performance diagrams of five methods in the first embodiment of the present invention in the independent test set;

[0030] Figure 5 In it, A - E are respectively the experimental diagrams of the influence of the training set scale of the cell line A549, cell line HCT116, cell line HepG2, cell line K562, and cell line MCF-7 on the performance of different prediction models in the first embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0031] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0032] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention.

[0033] In the case of no conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0034] Embodiment 1

[0035] This embodiment discloses an enhancer activity prediction method based on a bidirectional long short-term memory network, including:

[0036] Obtaining DNA sequence data and epigenetic modification data;

[0037] Encoding and processing the DNA sequence data respectively by reverse complementary k-mer, mismatched k-mer and word vector methods;

[0038] Concatenating the sequence features encoded by reverse complementary k-mer and mismatched k-mer in dimension to obtain concatenated features, and using the concatenated features as the input of the sequence extraction module;

[0039] Based on the context information of the sequence features encoded by the word vector method, obtaining a feature representation, and using the feature representation as the input of the Word2Vec module;

[0040] Using the epigenetic modification data as the input of the epigenetic modification module;

[0041] Fusing the output results of the sequence extraction module, the Word2Vec module and the epigenetic modification module, and determining the activity of the enhancer according to the fusion result.

[0042] In this embodiment, a prediction model, namely EAP-LSTM, is constructed to realize the prediction of the activity of enhancers in DNA fragments, and the prediction model is trained using a benchmark dataset, which includes two species, Drosophila and human.

[0043] The Drosophila dataset is generated by UMI-STARR-seq and covers the activities of 11,658 development-related and 7,062 housekeeping gene enhancers in the Drosophila S2 cell genome. The human dataset is provided by ENCODE and includes data from five different cell lines, namely A549, HCT116, HepG2, K562 and MCF-7.

[0044] It consists of enhancer activity data derived from STARR-seq, histone modification data obtained by ChIP-seq analysis, and chromatin accessibility data obtained by DNA-seq.

[0045] For the positive and negative samples of the Drosophila dataset, this embodiment directly uses the data provided by previous studies on github. For the human dataset, this embodiment downloads the "bed" file from the ENCODE database, then uses GKMSVM to extract sequences from the "bed" file, extends 500bp on each side of the peak to generate positive samples, and obtains 1001bp sequences; for negative samples, sequences with a GC content similar to that of the positive samples are selected as negative samples. For enhancer activity signals and epigenomic genetic data, the deepTools tool is used to extract the signals of positive and negative samples from the bigwig file, and the data is normalized using the log2(1 + signal) transformation to reduce bias.

[0046] The data obtained above is encoded by reverse complementary k-mer (RCkmer), mismatch k-mer, and Word2Vec word vector technology respectively. RCkmer encoding captures the information of the reverse complementary strand in the DNA sequence, enhancing the integrity of sequence expression; mismatch k-mer encoding improves the robustness of the model to genomic variations and minor differences by tolerating sequence variations; while Word2Vec automatically extracts the semantic features of the sequence by learning the context relationship in the DNA sequence, providing a richer representation and effectively improving the generalization ability of the model on large-scale datasets. The combination of these three encoding methods can more comprehensively capture the key information in the DNA sequence, thereby improving the prediction accuracy and reliability of genomics tasks.

[0047] Reverse complementary k-mer (RCkmer) is obtained by reversing the order of nucleotides and then taking the complement of each nucleotide for the sequence after k-mer encoding operation. k-mer is a nucleotide sequence of length 'k' in a larger nucleic acid sequence. For DNA, the complement of adenine (A) is thymine (T), and the complement of cytosine (C) is guanine (G), and vice versa. RCkmer is a variant of traditional k-mer, in which k-mer is not specific to a certain strand, so the reverse complement sequences are combined into a single feature. For example, 'AA' is the reverse complement of 'TT', 'GCA' is the reverse complement of 'TGC', and 'ACGTA' is the reverse complement of 'TACGT'. In this embodiment, the length of k-mer, i.e., k = 3, is set, so the feature dimension based on RCkmer is 32.

[0048] A mismatch k-mer is also a variant of the standard k-mer. When calculating the occurrence frequency of a certain k-mer, not only the original sequence of the k-mer is considered, but also the occurrence frequencies of all its possible mismatch variants (i.e., sequences formed by changing the base at a certain position in the k-mer). For example, when calculating the occurrence times of a k-mer with a sequence length of three ('GCA'), it is necessary to consider the original sequence ('GCA'), the change of the base at the first position of the k-mer ('ACA', 'CCA', 'TCA'), the change of the base at the second position of the k-mer ('GAA', 'GGA', 'GTA'), and the change of the base at the third position of the k-mer ('GCC', 'GCG', and 'GCT'). The sum of the occurrence times of these 3-tuples is regarded as the occurrence times of 'GCA'. In this embodiment, k can be set to 3, so the feature dimension based on mismatch is 64.

[0049] Word2Vec word vector technology has achieved great success in the field of natural language processing (NLP). Recently, these technologies have been widely applied in the bioinformatics community to address a significant limitation: even if the order of different sequences is reversed, k-mer-based features may exhibit a high degree of similarity. This adoption aims to overcome the limitations brought by the inherent limitations of k-mer feature representation in sequence analysis.

[0050] This embodiment adopts the idea of word vectors, takes all DNA sequences in the enhancer dataset as the corpus, considers each DNA sequence as a sentence in the corpus, and uses the 4 k types to form the vocabulary, considering each k-mer fragment as a word in the vocabulary. To ensure the complete independence of the independent test set, this embodiment only uses the sequences in the training set to form the corpus and the vocabulary. The continuous skip-gram model in Word2Vec is used to learn the feature vectors of each "word", and these word vectors can be used to train the language model subsequently; then the feature vectors of all words in the sequence are concatenated and used as the features of the sequence. Assuming that each "word" is embedded as a feature vector with a dimension of D and the sequence length is L, the feature dimension of each sequence is D×(L - k + 1). In this embodiment, the length of each word is set to 3 and is embedded as a 100-dimensional feature vector. Therefore, the feature dimension based on Word2Vec is 100×(L - 2).

[0051] Such as Figure 2As shown, the DNA fragments extracted from the dataset are encoded by three encoding methods: RCkmer, mismatch, and Word2Vec, respectively. Subsequently, the encoded sequence fragments are input into the prediction model for training, so that the prediction model can quantitatively predict the enhancer activity of the corresponding sequence.

[0052] In this embodiment, the constructed prediction model includes three parts: a Word2Vec module, an epigenetic modification module, and a sequence extraction module. The combination of these modules not only improves the robustness of the model but also reduces the risk of overfitting.

[0053] In the Word2Vec module and the epigenetic modification module, four convolutional blocks connected in sequence are respectively adopted. Each convolutional block includes a combination of a one-dimensional convolutional layer and a corresponding one-dimensional max-pooling layer connected in sequence, and then the final output results of the four convolutional blocks are input into a Bidirectional LSTM layer. Bi-LSTM can make full use of the context information of the sequence, capture bidirectional dependencies, and provide richer feature representations.

[0054] In the sequence extraction module, a convolutional block is used, which includes a combination of a one-dimensional convolutional layer and a corresponding one-dimensional max-pooling layer connected in sequence. The convolutional layer captures complex features from the input through convolutional calculations, while the max-pooling layer realizes downsampling by selecting the maximum value of each sub-region.

[0055] Specifically, the convolutional layer defined in this embodiment has 64 filters, a stride of 1, and different kernel sizes for different layers. In addition, multiple max-pooling layers with a pooling size of 3 are constructed. To further enhance the generalization ability of the model and prevent overfitting, a dropout layer with a probability of 0.2 is introduced after the max-pooling layer in the convolutional block. This layer randomly removes some neural network units during the training phase, thereby improving the overall robustness of the model. In addition, the model uses ReLU as the activation function to add a non-linear component to the model and enhance the expressive ability of the model.

[0056] The Word2Vec module and the epigenetic modification module are calculated through two fully connected layers containing 256 neurons. At the same time, the sequence extraction module is calculated through a fully connected layer containing 100 neurons, and the "ReLU" activation function is selected for use in the three modules. Finally, EAP-LSTM combines the results of multiple modules and generates a prediction value through a fully connected layer containing a single neuron and a "Linear" activation function.

[0057] The output results of the Word2Vec module, the epigenetic modification module, and the sequence extraction module are input into the feature fusion module, where they are fused through flattening and concatenation operations. After being processed by a fully connected layer based on the output of the feature fusion module, the enhancer activity value of Zeng Qiangqiang is obtained.

[0058] This embodiment proposes a new deep learning framework EAP-LSTM, which adopts multi-module input, analyzes it from multiple angles and extracts features through operations such as convolution and pooling, and then quantitatively predicts enhancer activity based on the features. Compared with previous methods, the classification effect of this embodiment is very excellent and superior to other methods.

[0059] To verify whether the combined module can effectively improve the performance of the model, this embodiment uses three performance evaluation indicators (MSE, PCC, SCC) to evaluate the performance of the three modules respectively, and compares their performance with that of the combined module based on the independent test set of each cell line. Among them, MSE represents the error between the predicted value and the true value, and the lower the value, the better. PCC and SCC measure the similarity between the predicted value and the true value from different angles, and the higher the value, the better.

[0060] In Figure 3 ,"A" represents the sequence extraction module, "B" represents the epigenetic modification module, and "C" represents the Word2Vec module. As can be seen from Figure 3 EAP-LSTM can effectively improve the performance of the model by fusing the three modules. The performance of all three indicators has been improved on all cell lines. In particular, among the 5 cell lines in the human dataset, it is found that adding epigenetic information can effectively improve the enhancer activity prediction ability, which proves the potential relationship between epigenetic modification and enhancers.

[0061] To prove the superiority of the EAP-LSTM proposed in this embodiment in quantitatively predicting enhancer activity, in the human dataset, it is compared with two previously invented activity prediction frameworks DeepSTARR and HEAP. Since there is only epigenetic modification data in the human dataset, in order to ensure a fair and meaningful comparison, each method without adding the epigenetic modification module in the human dataset (EAP-LSTM-DNA and HEAP-DNA methods) is also evaluated.

[0062] The performance of these five methods on the independent test set ( Figure 4 A-B) is compared. Compared with the previous methods, EAP-LSTM performs excellently on two indicators of all cell lines, which shows that the performance of EAP-LSTM is significantly better than the previous two methods, proving the superiority and robustness of EAP-LSTM in quantitatively predicting enhancer activity.

[0063] To demonstrate the robustness of this embodiment, EAP-LSTM is applied to a small-sample dataset. By setting different training set sizes, which are 10%, 25%, 50%, and 75% of the original dataset respectively, the impact of different data volumes on the model performance is simulated. The experimental results are as Figure 5 shown. EAP-LSTM exhibits excellent performance under most dataset sizes, and its performance is significantly better than other comparison methods. This result indicates that EAP-LSTM has stronger generalization ability when dealing with small-sample data and can achieve more accurate predictions under limited data conditions. As the amount of training data increases, the performance of each method generally improves, but EAP-LSTM always maintains a leading position under all sizes, further verifying its robustness and advantages under various data scales.

[0064] Embodiment 2

[0065] The purpose of this embodiment is to provide an enhancer activity prediction system based on a bidirectional long short-term memory network, including:

[0066] An acquisition module, which is configured to: acquire DNA sequence data and epigenetic modification data;

[0067] An encoding module, which is configured to: perform encoding processing on the DNA sequence data respectively through reverse complementary k-mer, mismatched k-mer, and word vector methods;

[0068] A first extraction module, which is configured to: splice the sequence features encoded by reverse complementary k-mer and mismatched k-mer in terms of dimension to obtain spliced features, and perform feature extraction on the spliced features through a sequence extraction module;

[0069] A second extraction module, which is configured to: obtain a feature representation based on the context information of the sequence features encoded by the word vector method, and perform feature extraction on the feature representation through a Word2Vec module based on a bidirectional long short-term time memory network;

[0070] A third extraction module, which is configured to: perform feature extraction on the epigenetic modification data through an epigenetic modification module based on a bidirectional long short-term time memory network;

[0071] A prediction module, which is configured to: fuse the output results of the sequence extraction module, the Word2Vec module, and the epigenetic modification module, and determine the activity of the enhancer according to the fusion result.

[0072] In more embodiments, there is also provided:

[0073] An electronic device includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment 1 is completed. For the sake of brevity, it will not be elaborated here.

[0074] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0075] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0076] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the method described in Embodiment 1 is completed.

[0077] The method in Embodiment 1 can be directly embodied as being executed and completed by a hardware processor, or can be executed and completed by a combination of hardware and software modules in the processor. The software modules may be located in mature storage media in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0078] A computer program product includes a computer program. When the computer program is executed by the processor, the method described in Embodiment 1 is implemented and completed.

[0079] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to perform the process / method as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules can be combined or divided as needed. The machine-executable instructions for program modules can be executed locally or within a distributed device. In a distributed device, program modules can be located in local and remote storage media.

[0080] The computer program code for implementing the method of the present invention can be written in one or more programming languages. This computer program code can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program code is executed by the computer or other programmable data processing devices, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the computer, partially on the computer, as an independent software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.

[0081] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier so that the device, apparatus, or processor can execute the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals can include electrical, optical, radio, acoustic, or other forms of propagated signals, such as carrier waves, infrared signals, etc.

[0082] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0083] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A method for predicting enhancer activity based on a bidirectional long short-term memory network, characterized in that: include: Obtain DNA sequence data and epigenetic modification data; The DNA sequence data is encoded by reverse complementary k-mer, mismatched k-mer and word vector methods respectively; Dimensionally splicing the sequence features encoded by the reverse complementary k-mer and the mismatched k-mer to obtain spliced ​​features, and extracting the spliced ​​features through a sequence extraction module; Based on the context information of the sequence features encoded by the word vector method, a feature representation is obtained; the feature representation and the epigenetic modification data are subjected to feature extraction through a Word2Vec module based on a bidirectional long short-term memory network and an epigenetic modification module respectively; The output results of the sequence extraction module, the Word2Vec module and the epigenetic modification module are fused, and the activity of the enhancer is determined according to the fusion results.

2. The enhancer activity prediction method based on bidirectional long short-term memory network according to claim 1, characterized in that: The DNA sequence data is coded by reverse complementary k-mer, specifically: Performing k-mer encoding operation on the DNA sequence data; The k-mer encoded sequence is obtained by reversing the order of nucleotides and taking the complement of each nucleotide.

3. The enhancer activity prediction method based on bidirectional long short-term memory network according to claim 1, characterized in that: The DNA sequence data is coded by mismatched k-mers, specifically: when calculating the occurrence frequency of a k-mer, not only the original sequence of the k-mer is considered, but also the occurrence frequencies of all possible mismatched variants of the k-mer.

4. The method for predicting enhancer activity based on a bidirectional long short-term memory network according to claim 1, characterized in that: The DNA sequence data is encoded using a word vector method, specifically: Treat each DNA sequence as a sentence; Using 4 of the k-mer fragment k Type forms a vocabulary, treating each k-mer segment as a word in the vocabulary; where k is the length of the k-mer; Obtain the feature vector of each word through the continuous skipping model in Word2Vec; The acquired feature vectors of each word are concatenated, and the concatenated feature vectors are used as the feature vectors of the DNA sequence.

5. The method for predicting enhancer activity based on a bidirectional long short-term memory network according to claim 1, characterized in that: The sequence features encoded by reverse complementary k-mer and mismatched k-mer are spliced ​​to obtain spliced ​​features; specifically: The sequence features after reverse complementary k-mer and mismatched k-mer encoding are subjected to convolution processing respectively; Flatten the sequence features after convolution processing; The sequence features processed by the flattening operation are concatenated through a fully connected layer to obtain concatenated features.

6. The method for predicting enhancer activity based on a bidirectional long short-term memory network according to claim 1, characterized in that: The processing process of the Word2Vec module based on the bidirectional long short-term memory network on the feature representation is: Performing convolution processing on the feature representation using a plurality of convolution blocks connected in sequence; The results processed by the convolutional block are input into the bidirectional long short-term memory network to utilize the contextual information of the sequence and capture the bidirectional dependency.

7. Enhancer activity prediction system based on bidirectional long short-term memory network, characterized in that: include: An acquisition module, which is configured to: acquire DNA sequence data and epigenetic modification data; An encoding module is configured to: encode the DNA sequence data by using reverse complementary k-mer, mismatched k-mer and word vector methods respectively; The first extraction module is configured to: splice the sequence features encoded by the reverse complementary k-mer and the mismatched k-mer in dimensions to obtain spliced ​​features, and extract features from the spliced ​​features through the sequence extraction module; The second extraction module is configured to: obtain a feature representation based on the context information of the sequence feature encoded by the word vector method; extract the feature representation and the epigenetic modification data through a Word2Vec module based on a bidirectional long short-term memory network and an epigenetic modification module respectively; The prediction module is configured to: fuse the output results of the sequence extraction module, the Word2Vec module and the epigenetic modification module, and determine the activity of the enhancer according to the fusion result.

8. An electronic device, characterized in that: The method comprises a memory and a processor and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method according to any one of claims 1 to 6 is completed.

9. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, complete the method described in any one of claims 1 to 6.

10. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.