Plant transcription factor binding site prediction method fusing complete dynamic convolution and CubeMLP
By fusing the method of complete dynamic convolution and CubeMLP, the DNA sequence and spatial structural characteristics are extracted and fused, and the problem of inaccuracy and inefficiency of predicting plant transcription factor binding sites in the prior art is solved, achieving high accuracy and stable prediction effects.
Patent Information
- Application Number
- CN202510066044.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-01-16
AI Technical Summary
Existing deep learning methods have problems of accuracy and inefficiency in plant transcription factor binding sites (TFBS) prediction, especially in the recognition of multi-class TFBS.
Using the method of fusion of complete dynamic convolution and CubeMLP, the DNA sequence data is subjected to One-hot encoding and DNA spatial structure feature extraction, combined with multi-dimensional attention mechanism and convolutional neural network, the multi-modal features of DNA sequence and spatial structure features are captured, and the features are efficiently fusion is achieved through CubeMLP.
The accuracy and efficiency of predicting binding sites of plant transcription factor are improved, and effective identification and classification of the characteristics of binding sites of different categories of transcription factor are achieved, with high prediction accuracy and stability.
Smart Images

Figure CN120148607A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics technology, and in particular relates to a method for predicting plant transcription factor binding sites by integrating full dynamic convolution with CubeMLP. Background Art
[0002] The transcription process is closely related to gene expression, and transcription factors (TF) play a vital role in plant growth, development, and environmental adaptability. TF is closely related to plant embryonic development, organ differentiation, and other morphological processes, and can regulate the biosynthesis of various metabolites in plants, which has an important impact on plant disease resistance, insect resistance, and stress resistance. TF affects plant transcription and gene expression by binding to specific DNA sequences (i.e., TFBS). Therefore, accurate prediction of TFBS is of great value for studying the TF-DNA binding mechanism and understanding the transcriptional regulation mechanism of plants.
[0003] In recent years, with the continuous advancement of machine learning and deep learning technologies, they have been successfully applied to computer vision, natural language processing and other fields. Deep learning methods can automatically learn features from large-scale raw data and mine high-order complex nonlinear information. Therefore, applying deep learning to the prediction of plant TFBS is expected to improve the accuracy and efficiency of prediction.
[0004] However, the current plant TFBS prediction methods based on deep learning are still in their infancy and have not yet formed a complete technical system. On the one hand, the existing methods only use DNA sequence information to predict TFBS. The representative deep learning methods are TSPTFBS1.0 and TSPTFBS2.0, which show strong robustness in the identification of single-class TFBS, but are still challenging in the identification of multiple-class TFBS. On the other hand, it is necessary to consider the complex binding process of plant TFBS and use the sequence information of DNA and the spatial structural characteristics of the DNA double helix to improve the accuracy of TFBS prediction. For example, the PlantBind method can realize the identification of multiple types of TFBS. However, this method does not fully integrate the characteristics of different modalities, resulting in certain limitations in the recognition accuracy. Therefore, it is of great research significance to develop a method for predicting plant transcription factor binding sites with high accuracy. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide a plant transcription factor binding site prediction method integrating full dynamic convolution and CubeMLP, which is used to predict transcription factor binding sites.
[0006] The technical solution adopted by the present invention to solve the above technical problems is: a method for predicting plant transcription factor binding sites by integrating full dynamic convolution and CubeMLP, comprising the following steps: S1: Divide the DNA sequence data and perform one-hot encoding to obtain a one-hot encoding matrix; use the Monte Carlo simulation method to extract the DNA spatial structure features to obtain a DNA spatial structure encoding matrix; S2: For the one-hot encoding matrix, adaptively adjust the convolution kernel weights through the multi-dimensional attention mechanism of the fully dynamic convolution network, and extract the DNA sequence features according to the characteristics of each type of transcription factor binding site sequence; S3: For the DNA spatial structure encoding matrix, utilize the weight sharing characteristic of the convolution operation of the convolutional neural network to extract the DNA spatial structure features; S4: Use CubeMLP to capture the complex spatial interaction relationships of cross-modal features through the MLP mechanisms in the sequence, modality, and channel dimensions, and fuse the multi-modal features of DNA sequence features and DNA spatial structure features; S5: Map the result of the fused features to different types of transcription factor binding sites to achieve the classification of the features of different transcription factor binding sites.
[0007] According to the above scheme, before step S1, the following steps are further included: S0: Sequence the biological material to obtain the DNA sequence data of the transcription factor binding site region of the plant.
[0008] According to the above scheme, in step S1, the specific steps are: Divide the DNA sequence data according to a fixed length and perform one-hot encoding to obtain a one-hot encoding matrix; Use the Monte Carlo simulation method to extract 14 different structural features including the features within 6 bases of each base in the DNA sequence, the features of 6 base pairs in the double helix structure, 1 minor groove width, and 1 minor groove potential to obtain a DNA spatial structure encoding matrix.
[0009] According to the above scheme, in step S2, the specific steps are: S21: Through the Avgpooling layer and 1×1 convolution, compress the overall features of the TF sequence in the sequence dimension and channel dimension of the one-hot encoding matrix respectively to obtain a compressed and pooled feature vector; S22: Extract different types of attention through the multi-head attention module, including convolution kernel attention, channel attention, filter attention, and spatial attention; S23: Weight the convolution layer weights one by one in the order of space, channel, filter, and convolution kernel to obtain the convolution layer weights after full-dimensional weighting; S24: Use the weighted dynamic convolution layer to perform convolution operations on the one-hot encoding matrix to gradually extract local features; S25: Rapidly downsample the local features through a downsampling layer, and use a linear layer to map them to high-order features of a specified dimension.
[0010] According to the above solution, in step S3, the specific steps are as follows: S31: Gradually extract local features from the DNA spatial structure encoding matrix through a convolutional block; S32: Rapidly downsample the local features through a downsampling layer, and use a linear layer to map them to high-order features of a specified dimension.
[0011] According to the above solution, in step S4, the specific steps are as follows: S41: Concatenate the DNA sequence features and the DNA spatial structure features to obtain multimodal fusion data; S42: Use CubeMLP to process the multimodal fusion data to obtain fully fused features.
[0012] According to the above solution, in step S5, the specific steps are as follows: S51: Comprehensively fuse the features through a fully connected layer; S52: Use the activation function Sigmoid to obtain the probability of the transcription factor binding site features, and realize the classification of different transcription factor binding site features.
[0013] According to the above solution, the following steps are also included: S6: Use the binary cross-entropy loss function BCELoss that measures the difference between the target label value and the predicted probability value to calculate the network loss; use the optimized gradient algorithm AdamW to update the parameters, and add a weight decay term to the loss function to adjust the parameters during the adaptive learning rate update process.
[0014] A DeepTFBS model includes an input layer, a feature extraction layer, a cube multi-layer perceptron layer, and an output layer; The input layer includes a one-hot encoding layer and a spatial structure encoding layer; the one-hot encoding layer is used to perform data partitioning and one-hot encoding on the DNA sequence data to obtain a one-hot encoding matrix; the spatial structure encoding layer is used to extract DNA spatial structure features by using the Monte Carlo simulation method to obtain a DNA spatial structure encoding matrix; The feature extraction layer includes a fully dynamic convolutional network and a convolutional neural network; the fully dynamic convolutional network is used to adaptively adjust the convolutional kernel weights for the one-hot encoding matrix through a multi-dimensional attention mechanism, and extract DNA sequence features according to the characteristics of each type of transcription factor binding site sequence; the convolutional neural network is used to extract DNA spatial structure features by using the weight sharing characteristic of the convolutional operation for the DNA spatial structure encoding matrix; The cube multi-layer perception layer, including a sequence-dimension multi-layer perceptron, a modality-dimension multi-layer perceptron, and a channel-dimension multi-layer perceptron, is used to capture the complex spatial interaction relationships of cross-modal features through the MLP mechanisms in the three dimensions of sequence, modality, and channel, and fuse the multi-modal features of DNA sequence features and DNA spatial structure features; The output layer, including several multi-layer perceptrons, is used to map the result of the fused features into different categories of transcription factor binding sites, and achieve the classification of the features of different transcription factor binding sites.
[0015] Furthermore, the fully dynamic convolutional network includes two dynamic convolutional blocks, a downsampling layer, and a linear layer; the dynamic convolutional block includes a one-dimensional dynamic convolution, a RELU activation function, and a Dropout layer; the dynamic convolution includes a pooling-excitation multi-head attention module for realizing the adaptive adjustment of the convolutional kernel weights; The convolutional neural network includes two convolutional blocks, a downsampling layer, and a linear layer; the convolutional block includes a one-dimensional convolution, a RELU activation function, and a Dropout layer; the convolution has the characteristics of weight sharing and local receptive fields for extracting the features of the DNA spatial structure with fewer parameters; the RELU activation function and the Dropout layer enhance the generalization ability of the model by introducing non-linear features.
[0016] The beneficial effects of the present invention are as follows: 1. The method for predicting plant transcription factor binding sites by fusing the fully dynamic convolution and CubeMLP of the present invention performs One-hot encoding processing on the DNA sequence data of the plant transcription factor binding site region obtained by sequencing and extracts the spatial structure features of the DNA double helix; uses the ODConv network to dynamically extract the pattern features of different categories of TFBS sequences, uses the convolutional neural network to extract the DNA spatial structure features, and introduces non-linear features through the RELU activation function and the Dropout layer, thereby enhancing the generalization ability of the model; uses the idea of the multi-axis MLP of CubeMLP to achieve the fusion of multi-axis complex features through the three dimensions of sequence, modality, and channel, and realizes the efficient fusion of the DNA sequence feature modality and the DNA spatial structure modality feature; uses the fully connected layer and the Sigmoid function to map the result of the fused features into different categories of transcription factor binding sites, and further realizes the recognition of different categories of transcription factor binding sites; realizes the function of predicting transcription factor binding sites.
[0017] 2. The present invention identifies different categories of transcription factor binding sites by learning the features of plant DNA sequences and spatial structures, has a high prediction accuracy, provides a new reference for the research on the recognition sites of plant transcription factors, and has a wide application prospect.
[0018] 3. Experimental results of predicting 315 types of TFBS in Arabidopsis thaliana by the present invention show that, compared with existing models, the DeepTFBS method of the present invention has the highest prediction accuracy; in terms of predicting TFBS by chromosome, the prediction accuracy of DeepTFBS is better than existing methods, indicating that DeepTFBS has strong stability; ablation experiment results of the ODConv and CubeMLP modules in the DeepTFBS method show that both ODConv and CubeMLP make great contributions in prediction.
[0019] 4. The DeepTFBS method proposed by the present invention can predict all potential TFBS at the plant genome level, can supplement the data volume of TFBS and the TFBS obtained by CHIP-seq experiments, and has portability and generalizability.
[0020] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 It is a flowchart of an embodiment of the present invention.
[0023] Figure 2 It is a principle block diagram of an embodiment of the present invention.
[0024] Figure 3 It is a binary classification and multi-classification result diagram of predicting 315 types of TFBS in Arabidopsis thaliana by the DeepTFBS method of an embodiment of the present invention.
[0025] Figure 4 It is a multi-classification prediction result diagram of different methods for 315 types of TFBS in Arabidopsis thaliana.
[0026] Figure 5 It is a performance and stability result diagram of 5-fold cross-validation of chromosome-by-chromosome DeepTFBS in multi-classification prediction of 315 types of TFBS in Arabidopsis thaliana in an embodiment of the present invention.
[0027] Figure 6 It is an overlapping result diagram of 5 potential TFBS predicted by the DeepTFBS of an embodiment of the present invention for RAP211, BIM2, SPL5, CBF4, SPL9, SPL5 in Arabidopsis thaliana and the TFBS obtained by CHIP-seq sequencing. Detailed implementation manners
[0028] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0029] Embodiment 1 Refer to Figure 1 , and the specific steps of the method for predicting plant transcription factor binding sites by integrating fully dynamic convolution and CubeMLP are as follows: S1: Perform data partitioning and One-hot encoding on the DNA sequence data to obtain a One-hot encoding matrix; adopt the Monte Carlo simulation method to extract the DNA spatial structure features to obtain a DNA spatial structure encoding matrix; S2: For the One-hot encoding matrix, adaptively adjust the convolution kernel weights through the multi-dimensional attention mechanism of the fully dynamic convolution network, and extract the DNA sequence features according to the characteristics of each type of transcription factor binding site sequence; S3: For the DNA spatial structure encoding matrix, utilize the weight sharing characteristic of the convolution operation of the convolutional neural network to extract the DNA spatial structure features; S4: Use CubeMLP to capture the complex spatial interaction relationships of cross-modal features through the MLP mechanisms in the three dimensions of sequence, modality, and channel, and fuse the multi-modal features of DNA sequence features and DNA spatial structure features; S5: Map the result of the fused features to different types of transcription factor binding sites to realize the classification of the features of different transcription factor binding sites.
[0030] Before step S1, the following steps are further included: S0: Sequence the biological material to obtain the DNA sequence data of the plant transcription factor binding site region.
[0031] According to the above solution, in step S1, the specific steps are: Partition the DNA sequence data by a fixed length and perform One-hot encoding to obtain a One-hot encoding matrix; Use the Monte Carlo simulation method to extract 14 different structural features including the features within 6 bases of each base in the DNA sequence, the features of 6 base pairs in the double helix structure, 1 minor groove width, and 1 minor groove potential to obtain a DNA spatial structure encoding matrix.
[0032] Further, in step S2, the specific steps are: S21: Overall feature compression of the TF sequence is performed on the One-hot encoded matrix in the sequence dimension and channel dimension respectively through the Avgpooling layer and 1×1 convolution to obtain the compressed and pooled feature vector; S22: Different types of attention are extracted through the multi-head attention module, including convolutional kernel attention, channel attention, filter attention, and spatial attention; S23: The weights of the convolutional layer are weighted one by one in the order of space, channel, filter, and convolutional kernel to obtain the fully dimension-weighted convolutional layer weights; S24: The weighted dynamic convolutional layer is used to perform convolution operations on the One-hot encoded matrix to gradually extract local features; S25: The local features are rapidly downsampled through the downsampling layer, and a linear layer is used to map them to high-order features of the specified dimension.
[0033] In step S3, the specific steps are as follows: S31: Local features are gradually extracted from the DNA spatial structure encoding matrix through the convolutional block; S32: The local features are rapidly downsampled through the downsampling layer, and a linear layer is used to map them to high-order features of the specified dimension.
[0034] In step S4, the specific steps are as follows: S41: The DNA sequence features and DNA spatial structure features are concatenated to obtain the multimodal fusion data; S42: The multimodal fusion data is processed using CubeMLP to obtain the fully fused features.
[0035] In step S5, the specific steps are as follows: S51: The features are comprehensively fused through the fully connected layer; S52: The probability of the transcription factor binding site features is obtained using the activation function Sigmoid to achieve the classification of different transcription factor binding site features.
[0036] The following steps are also included: S6: The binary cross-entropy loss function BCELoss that measures the difference between the target label value and the predicted probability value is used to calculate the network loss; the parameters are updated using the optimized gradient algorithm AdamW, and the weight decay term is added to the loss function to adjust the parameters during the adaptive learning rate update process.
[0037] In this embodiment, the DNA sequence data of the plant transcription factor binding site region obtained by sequencing is processed by One-hot encoding and the spatial structure features of the DNA double helix are extracted; the ODConv network is used to dynamically extract the pattern features of different categories of TFBS sequences, the convolutional neural network is used to extract the DNA spatial structure features, and the RELU activation function and the Dropout layer are used to introduce non-linear features, thereby enhancing the generalization ability of the model; the idea of the multi-axis MLP of CubeMLP is used to achieve the fusion of multi-axis complex features through three dimensions of sequence, modality, and channel, and the efficient fusion of DNA sequence feature modality and DNA spatial structure modality features is realized; the fully connected layer and the Sigmoid function are used to map the results of the fused features into different categories of transcription factor binding sites, and then the recognition of different categories of transcription factor binding sites is realized; the function of predicting transcription factor binding sites is realized.
[0038] Example 2 The steps of this embodiment are the same as those of Embodiment 1, except that each step is applied to a specific example. Specifically, it includes the following steps: S0: Sequencing biological materials to obtain DNA sequence data of the plant transcription factor binding site region; S1: Performing data partitioning and One-hot encoding on the DNA sequence data to obtain a One-hot encoding matrix; using the Monte Carlo simulation method to extract the DNA spatial structure features to obtain a DNA spatial structure encoding matrix; Performing data partitioning and other processing on the data of the plant transcription factor binding site TFBS region obtained by sequencing, and using the original DNA sequence with a length of 101bp S ={ s 1, s 2... s i}, s i ∈{A, T, C, G} as the input of the model. Then, the sequence is One-hot encoded, the base A is converted to [1, 0, 0, 0], the base T is converted to [0, 1, 0, 0], the base C is converted to [0, 0, 1, 0], and the base G is converted to [0, 0, 0, 1], and then a 4*101 One-hot encoding matrix Seq ( Seq ∈ R 4×101 ) is obtained, and this matrix is input into the fully dynamic convolutional network (ODConv) for the next step of processing.
[0039] Then, the Monte Carlo simulation method is used to extract 14 different structural features composed of 6 base-internal features (Buckled, Sheared, Stretched, Propeller Twisted (ProT), Opened, and Staggered) for each base in the sequence, 6 base-pair features (Tilted, Shifted, Slid, Rolled, Risen, and Helix Twisted) in the double helix structure, 1 minor groove width, and 1 minor groove potential. Finally, a DNA spatial structure encoding matrix is obtained. Shape ( Shape ∈R 14×101 ), and it is input into a convolutional neural network for the next step of processing.
[0040] S2: Use the ODConv network to extract unique patterns and features of multiple types of transcription factor binding site (TFBS) sequences. Dynamic convolution enhances the feature extraction ability of the convolutional layer by introducing a multi-dimensional attention mechanism. This operation allows the convolutional kernel weights to be adaptively adjusted, enabling each type of TFBS sequence to be processed with different convolutional kernels, so that the model can extract the best feature representations according to the characteristics of each type of TFBS sequence.
[0041] The ODConv network consists of two dynamic convolution blocks, a downsampling layer, and a linear layer. Among them, the dynamic convolution block contains a one-dimensional dynamic convolution, a RELU activation function, and a Dropout layer. The adaptive weights of the dynamic convolution are realized by a pooling-excitation multi-head attention module to adaptively adjust the convolutional kernel weights.
[0042] S21: The input sequence Seq is processed through an Avgpooling layer and a 1×1 convolution. This operation realizes the overall feature compression of the TF sequence in the sequence dimension and the channel dimension, and obtains the compressed and pooled feature vector Seq G ( Seq G ∈R C×1 ). Specifically, it is shown in formula (1).
[0043] (1) S22: Extract different types of attention through the multi-head attention module, including convolutional kernel attention a k ( a k ∈R K num), channel attention a c ( a c ∈RC in), filter attention a f ( a f ∈R C out), spatial attention a s ( a s ∈R K size), and the specific implementation process is shown in Formulas (1) - (5).
[0044] (2) (3) (4) (5) S23: Weight the convolutional layer weights one by one in the order of space, channel, filter, and convolution kernel to obtain the convolutional layer weights after full - dimension weighting W ODconv ( W ODconv ∈R C out ×C in ×K size), ensuring that the convolution operation is different in all dimensions. The specific implementation is shown in Formula (6).
[0045] (6) Among them, K num is the number of convolution kernels, C out is the number of filters in each convolution kernel, W i j ( W i j ∈R Cin ×Ksize ) represents the i th filter of theth convolution kernel. j a s , a c are the attentions assigned to the convolution kernel in the spatial dimension and channel dimension respectively, ai f is the attention assigned to different filters of the convolution kernel, a k and
[0046] S24: Use the weighted dynamic convolutional layer for SeqPerform a convolution operation, and also use the ReLU activation function and Dropout to introduce non-linear features. The complete implementation process of the dynamic convolution block is shown in Equation (7).
[0047] (7) The ODConv network uses ODConv 1 and ODConv 2 two convolution blocks to gradually extract local features from the encoded Seq matrix Seq '' . The specific implementation is shown in Equation (8).
[0048] (8) S25: Downsample the local features Seq '' quickly through the downsampling layer, and use the linear layer to map it to high-order features of the specified dimension Seq * . The specific implementation is shown in Equation (9).
[0049] (9) S3: Utilize the weight sharing strategy of the convolutional neural network to effectively extract the spatial structure features of DNA with fewer parameters. In addition, the ReLU activation function and the Dropout layer introduce non-linear features, thus enhancing the generalization ability of the model.
[0050] The convolutional neural network layer consists of two convolutional blocks, a downsampling layer and a linear layer. The convolutional block contains a one-dimensional convolution, a ReLU activation function and a Dropout layer. The characteristics of weight sharing and local receptive fields in the convolution operation enable the effective extraction of the spatial structure features of DNA while using fewer parameters. The ReLU activation function and the Dropout layer introduce non-linear features, thus enhancing the generalization ability of the model. The specific implementation of the convolutional block is shown in Equation (10).
[0051] (10) S31: The convolution module uses Conv 1 and Conv 2 two convolutional blocks to gradually extract local features from the encoded Shape matrix Shape'' . The specific implementation is shown in Equation (11).
[0052] (11) S32: Downsample the local features through a downsampling layer Shape'' to quickly reduce the sampling rate, and then use a linear layer to map them to high-order features of a specified dimension Shape* . The specific implementation is shown in Equation (12).
[0053] (12) S4: To effectively fuse the extracted DNA sequence features and different modal features of the DNA spatial structure, CubeMLP is used to capture the complex spatial interaction relationships of cross-modal features, thereby realizing the fusion between multi-modal features. The CubeMLP structure uses MLP units in three dimensions of sequence (L), modality (M), and channel (C) to achieve the fusion of multi-axis complex features, ensuring the effective sharing of modal information in multiple dimensions while significantly reducing the computational cost.
[0054] S41: Concatenate the sequence features Seq* and the spatial structure features Shape* to obtain multi-modal fusion data S ( S ∈R L ×M×C ), and the specific implementation is shown in Equation (13).
[0055] (13) S42: Process the multi-modal fusion data using CubeMLP S . Each MLP unit consists of two fully connected layers, a GELU activation function, and a layer normalization. The GELU activation function based on probabilistic activation is introduced to capture more complex patterns. At the same time, during the feature fusion process, residual connections are used to prevent information loss in deep networks and accelerate model convergence. Specifically, the feature fusion in the sequence (L) dimension, modality (M) dimension, and channel (C) dimension is specifically implemented as shown in Equations (14) - (16).
[0056] (14) (15) (16) The CubeMLP layer uses two layers of CubeMLP stacked, and finally obtains the fully fused features S * ( S* ∈R L'×M'×C' ). Specifically, it is shown in Equation (17).
[0057] (17) S5: The output layer is used to predict the categories of TFBSs, achieving the classification of the characteristics of plant transcription factor binding sites.
[0058] S51: The output features of CubeMLP S * Input and output layer; integrated through two fully connected layers, and the fully connected neural network is used to further extract the fused features, capturing the hierarchical features of different transcription factors. The ReLU activation function and Dropout layer are used between the fully connected layers to further prevent overfitting; S52: Use the activation function Sigmoid to obtain the probability of transcription factor binding site features, and the value range is between 0 and 1. The specific implementation is shown in formula (18).
[0059] (18) S6: Calculate the network loss using the binary cross-entropy loss function BCELoss that measures the difference between the target label value and the predicted probability value, and update the parameters using the optimized gradient algorithm AdamW, which directly adds the weight decay term to the loss function to ensure that the parameters can be adjusted more accurately during the adaptive learning rate update process.
[0060] When calculating the network loss, use the Binary cross-entropy loss function BCELoss that measures the difference between the target label value and the predicted probability value. The specific implementation is shown in formula (19).
[0061] (19) p is the theoretical label, taking 0 or 1. q is the predicted value of the model output, and the value range is [0, 1], w is the weight.
[0062] Update the parameters using the optimized gradient algorithm AdamW, which directly adds the weight decay term to the loss function to ensure that the parameters can be adjusted more accurately during the adaptive learning rate update process. The specific implementation is shown in formula (20).
[0063] (20) Among them, weight_decay is the weight decay coefficient, lr is the learning rate, m is the first moment estimate of the gradient, v is the second moment estimate of the gradient. eps is a small constant used to increase numerical stability.
[0064] The present invention identifies binding sites of different classes of transcription factors by learning the characteristics of plant DNA sequences and spatial structures, with high prediction accuracy, providing a new reference for the research on the recognition sites of plant transcription factors and having broad application prospects.
[0065] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0066] Embodiment 3 The following are the example results of TFBS prediction using the method of the present invention in the data of 315 classes of TFBS in Arabidopsis thaliana. The superiority of the method of the present invention in TFBS prediction is detailed through relevant experiments, thereby providing a reference for the prediction of plant TFBS.
[0067] First, the model was trained and tested on the Arabidopsis thaliana dataset to evaluate the ability of DeepTFBS to predict plant TFBS. The specific experimental results are as Figure 3 shown. Seven evaluation metrics including SN, SP, ACC, AUC, AP, and MCC were used to evaluate the results of DeepTFBS for binary classification and multi-classification. In addition, without distinguishing the names of the evaluation metrics for binary classification and multi-classification, the binary classification metrics were named: SN_b, SP_b, ACC_b, AUC_b, AP_b, MCC_b; and the multi-classification metrics were named: SN_m, SP_m, ACC_m, AUC_m, AP_m, MCC_m;. In the multi-classification task, the AP and MCC metrics were mainly used for method comparison. Because the AP and MCC metrics can more accurately reflect the multi-classification ability of the model in the case of data imbalance. Specifically, the average AUC_b of DeepTFBS was 0.9714, the average AUC_m was 0.9755, the average AP_b was 0.9736, the average AP_m was 0.5017, the average MCC_b was 0.9310, and the average MCC_m was 0.3139. In the multi-classification case, the AP_m and MCC_m metrics decreased. This may be related to the serious class imbalance problem in the Arabidopsis thaliana dataset. In addition, there are similar motif sequences within the family among the TFBS of different TFs in Arabidopsis thaliana, which leads to misjudgment of negative samples. The distribution of each metric is as Figure 3 (F) shows that there are few outliers in the metrics and the distribution is concentrated. The above results indicate that DeepTFBS has stable classification performance and can effectively distinguish positive and negative samples.
[0068] Furthermore, to more deeply evaluate the multi-classification ability of DeepTFBS on the Arabidopsis dataset, four cutting-edge methods for predicting plant TFBS were trained and tested, including PlantBind, DeepSTF, D_SSCA, and DenseNet. The specific experimental results are as Figure 4 shown. For the average AUC_m metric, the result of DeepTFBS was 0.9755, which was 0.07% higher than that of PlantBind, 0.51% higher than that of DeepSTF, 19.03% higher than that of D_SSCA, and 47.2% higher than that of DenseNet. For the imbalanced dataset, the average MCC_m metric was used for the experiment. The results showed that the MCC_m metric of DeepTFBS was 0.3139, which was 0.41% higher than that of PlantBind, 1.46% higher than that of DeepSTF, 20.06% higher than that of D_SSCA, and 31.08% higher than that of DenseNet. Meanwhile, the distribution of each multi-classification metric was analyzed, as Figure 4 shown in Figure B, and it was found that the result distribution of DeepTFBS was more concentrated than that of other methods. In summary, DeepTFBS demonstrated the best multi-classification performance on the Arabidopsis dataset.
[0069] To further verify the stability of DeepTFBS, 5 cross-validation experiments by chromosome were conducted. Specifically, for the 5 chromosomes of Arabidopsis, all TFBS data on one chromosome were selected as the test set each time, and the rest were used as the training set. The results are as Figure 5 shown. It can be seen that each multi-classification metric decreased slightly. The average AUC_m was 0.9719, the average AP_m was 0.4818, and the average MCC_m was 0.3039. Ridge plots B, C, and D revealed that there were no significant distribution differences among the results of each cross-validation. Based on these results, it can be concluded that DeepTFBS has the ability to efficiently and stably predict plant TFBS based on a lightweight model architecture.
[0070] Furthermore, to evaluate the influence of the ODconv and CubeMLP parts on the prediction ability of the DeepTFBS model, ablation experiments were conducted for analysis. The ablation experiment results of different modules in the DeepTFBS method are shown in Table 1.
[0071] Table 1 Ablation Experiment Results of DeepTFBS
[0072] In each experiment, the hyperparameters were kept consistent with DeepTFBS. In Experiment A, only ODconv was used for feature extraction without feature fusion. In Experiment B, ODconv was replaced with Conv with the same parameters. The results showed that after removing CubeMLP, the average AUC_m decreased by 3.84% and the average AP_m decreased by 11.847%. This indicates that the fusion of CubeMLP in three dimensions helps to deeply fuse sequence features and structural features, and the fused representation is more conducive to the classification of the model. However, although the performance of Model B is slightly better than that of DeepTFBS, the static structure of Conv limits the robustness of Model B, and its accuracy in subsequent cross-species prediction is not high. It can be seen that the dynamic structure of ODconv enables it to adapt to different inputs, thus providing stronger robustness for DeepTFBS.
[0073] Based on the good prediction accuracy of DeepTFBS on the Arabidopsis thaliana dataset, all potential TFBS were further predicted at the Arabidopsis thaliana genome level. First, a sliding window method was adopted to intercept sequences from the Arabidopsis thaliana whole-genome dataset. The sliding window was set to 101bp and the sliding step was set to 10bp, so as to obtain 11,894,312 DNA sequences. Based on this dataset, the DNA sequence fragments were input into the trained DeepTFBS model for prediction. Taking 0.5 as the threshold, the sequences with predicted values greater than 0.5 are potential TFBS. Finally, the number of predicted potential TFBS far exceeds the number of TFBS obtained by actual sequencing. This difference may be due to the fact that the data obtained by actual sequencing was obtained at a specific time and in a single cell type, while DeepTFBS predicted all potential TFBS at different times and in different tissues at the genome level. The specific prediction results are as Figure 6 shown.
[0074] The experimental results of predicting 315 types of TFBS in Arabidopsis thaliana by the present invention show that compared with the existing models, the DeepTFBS method of the present invention has the highest prediction accuracy; in terms of predicting TFBS by chromosome, the prediction accuracy of DeepTFBS is better than the existing methods, indicating that DeepTFBS has strong stability; the ablation experiment results of the ODConv and CubeMLP modules in the DeepTFBS method show that both ODConv and CubeMLP make great contributions in prediction.
[0075] The DeepTFBS method proposed by the present invention can predict all potential TFBS at the plant genome level, can supplement the data volume of TFBS and the TFBS obtained by CHIP-seq experiments, and has portability and generalizability.
[0076] Example 4 This embodiment is used to implement the principle of the above method embodiment to construct a DeepTFBS model, which includes an input layer, a feature extraction layer, a cube multi-layer perceptron layer, and an output layer; The input layer includes a one-hot encoding layer and a spatial structure encoding layer; the one-hot encoding layer is used to partition and perform one-hot encoding on DNA sequence data to obtain a one-hot encoding matrix; the spatial structure encoding layer is used to extract DNA spatial structure features by using the Monte Carlo simulation method to obtain a DNA spatial structure encoding matrix; The feature extraction layer includes a fully dynamic convolutional network and a convolutional neural network; the fully dynamic convolutional network is used to adaptively adjust the convolutional kernel weights of the one-hot encoding matrix through a multi-dimensional attention mechanism, and extract DNA sequence features according to the characteristics of each type of transcription factor binding site sequence; the convolutional neural network is used to extract DNA spatial structure features by using the weight sharing characteristic of convolutional operations on the DNA spatial structure encoding matrix; The cube multi-layer perceptron layer includes a sequence dimension multi-layer perceptron, a modality dimension multi-layer perceptron, and a channel dimension multi-layer perceptron, which are used to capture the complex spatial interaction relationships of cross-modal features through the MLP mechanism in three dimensions of sequence, modality, and channel, and fuse the multi-modal features of DNA sequence features and DNA spatial structure features; The output layer includes several multi-layer perceptrons, which are used to map the result of the fused features to different types of transcription factor binding sites to achieve the classification of different transcription factor binding site features.
[0077] Among them, the fully dynamic convolutional network includes two dynamic convolutional blocks, a downsampling layer, and a linear layer; the dynamic convolutional block includes a one-dimensional dynamic convolution, a RELU activation function, and a Dropout layer; the dynamic convolution includes a gather-excite multi-head attention module for realizing the adaptive adjustment of convolutional kernel weights; The convolutional neural network includes two convolutional blocks, a downsampling layer, and a linear layer; the convolutional block includes a one-dimensional convolution, a RELU activation function, and a Dropout layer; the convolution has the characteristics of weight sharing and local receptive field for extracting the features of DNA spatial structure with fewer parameters; the RELU activation function and the Dropout layer enhance the generalization ability of the model by introducing non-linear features.
[0078] It should be noted that according to the needs of implementation, each step / component described in this application can be split into more steps / components, or two or more steps / components or partial operations of steps / components can be combined into new steps / components to achieve the purpose of the present invention.
[0079] This embodiment further includes a processor, a communication interface, a memory, and a communication bus; wherein the processor, the communication interface, and the memory complete communication with each other through the communication bus; a computer program is stored in the memory, and when the program is executed by the processor, the processor executes the steps of the method for predicting plant transcription factor binding sites by integrating fully dynamic convolution and CubeMLP.
[0080] This embodiment also provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by the processor, the processor implements the method for predicting plant transcription factor binding sites by integrating fully dynamic convolution and CubeMLP.
[0081] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.
[0082] Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0083] This application is described with reference to the flowcharts of the method and computer program product according to Embodiment 1 of the present application and the block diagrams of the device (system) according to Embodiment 4. It should be understood that each flow or block in the flowchart or block diagram can be implemented by computer program instructions, and the combination of the flows or blocks in the flowchart or block diagram can also be implemented by computer program instructions.
[0084] These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a DeepTFBS model for implementing the functions specified in one Figure 1 flow or multiple flows or blocks Figure 1 block or multiple blocks.
[0085] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one Figure 1 flow or multiple flows or blocks Figure 1 block or multiple blocks.
[0086] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for implementing the steps of the method for predicting plant transcription factor binding sites that fully integrates dynamic convolution and CubeMLP in the process Figure 1 a process or multiple processes or boxes Figure 1 or multiple boxes of the method for predicting plant transcription factor binding sites that fully integrates dynamic convolution and CubeMLP specified in the boxes.
[0087] The above embodiments are only used to illustrate the design concept and characteristics of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design concepts disclosed in the present invention are within the protection scope of the present invention.
Claims
1. A plant transcription factor binding site prediction method integrating fully dynamic convolution and CubeMLP, characterized by: The following steps are involved: S1: Perform data partitioning and One-hot encoding on DNA sequence data to obtain a One-hot encoding matrix; use Monte Carlo simulation method to extract DNA spatial structure features to obtain a DNA spatial structure encoding matrix; S2: For the one-hot encoding matrix, the convolution kernel weights are adaptively adjusted through the multi-dimensional attention mechanism of the fully dynamic convolutional network to extract DNA sequence features according to the characteristics of each type of transcription factor binding site sequence; S3: For the DNA spatial structure encoding matrix, the weight sharing property of the convolution operation of the convolutional neural network is used to extract the DNA spatial structure features; S4: CubeMLP is used to capture the complex spatial interaction relationship of cross-modal features through the MLP mechanism of sequence, modality and channel dimensions, integrating the multimodal features of DNA sequence features and DNA spatial structure features; S5: Map the results of the fusion features to different categories of transcription factor binding sites to achieve classification of different transcription factor binding site features.
2. The method for predicting plant transcription factor binding sites by integrating full dynamic convolution and CubeMLP according to claim 1, characterized in that: Before step S1, the method further includes the following steps: S0: Sequence biological materials to obtain DNA sequence data of the transcription factor binding site region of the plant.
3. The method for predicting plant transcription factor binding sites by integrating full dynamic convolution and CubeMLP according to claim 1, characterized in that: In the step S1, the specific steps are: Divide the DNA sequence data into fixed lengths and perform one-hot encoding to obtain a one-hot encoding matrix; Monte Carlo simulation method was used to extract 14 different structural features consisting of 6 intra-base features of each base in the DNA sequence, 6 base pair features in the double helix structure, 1 minor groove width and 1 minor groove potential, and the DNA spatial structure coding matrix was obtained.
4. The method for predicting plant transcription factor binding sites by integrating full dynamic convolution and CubeMLP according to claim 1, characterized in that: In the step S2, the specific steps are: S21: The overall feature compression of the TF sequence is performed on the One-hot encoding matrix in the sequence dimension and channel dimension through the Avgpooling layer and 1×1 convolution to obtain the compressed and aggregated feature vector; S22: Extract different types of attention through multi-head attention modules, including convolution kernel attention, channel attention, filter attention and spatial attention; S23: weight the convolution layer weights one by one in the order of space, channel, filter and convolution kernel to obtain the convolution layer weights after full-dimensional weighting; S24: Use the weighted dynamic convolution layer to perform convolution operations on the one-hot encoding matrix to gradually extract local features; S25: The local features are quickly downsampled through the downsampling layer and mapped to high-order features of the specified dimension using the linear layer.
5. The method for predicting plant transcription factor binding sites by integrating full dynamic convolution and CubeMLP according to claim 1, characterized in that: In the step S3, the specific steps are: S31: local features are gradually extracted from the DNA spatial structure encoding matrix through convolution blocks; S32: The local features are quickly downsampled through the downsampling layer and mapped to high-order features of the specified dimension using the linear layer.
6. The method for predicting plant transcription factor binding sites by integrating full dynamic convolution and CubeMLP according to claim 1, characterized in that: In the step S4, the specific steps are: S41: splicing DNA sequence features and DNA spatial structure features to obtain multimodal fusion data; S42: Use CubeMLP to process multimodal fusion data to obtain fully fused features.
7. The method for predicting plant transcription factor binding sites by integrating full dynamic convolution and CubeMLP according to claim 1, characterized in that: In the step S5, the specific steps are: S51: comprehensive fusion features through fully connected layers; S52: Use the activation function Sigmoid to obtain the probability of transcription factor binding site features and realize the classification of different transcription factor binding site features.
8. The method for predicting plant transcription factor binding sites by integrating full dynamic convolution and CubeMLP according to claim 1, characterized in that: The following steps are also included: S6: The network loss is calculated using the binary cross entropy loss function BCELoss, which measures the difference between the target label value and the predicted probability value; the parameters are updated using the optimized gradient algorithm AdamW, and the weight decay term is added to the loss function to adjust the parameters during the adaptive learning rate update process.
9. A DeepTFBS model, characterized in that: It includes input layer, feature extraction layer, cube multi-layer perception layer and output layer; The input layer includes a one-hot encoding layer and a spatial structure encoding layer; the one-hot encoding layer is used to perform data division and one-hot encoding on the DNA sequence data to obtain a one-hot encoding matrix; the spatial structure encoding layer is used to extract DNA spatial structure features using a Monte Carlo simulation method to obtain a DNA spatial structure encoding matrix; Feature extraction layer, including fully dynamic convolutional network and convolutional neural network; the fully dynamic convolutional network is used to adaptively adjust the convolution kernel weights of the one-hot encoding matrix through a multi-dimensional attention mechanism, and extract DNA sequence features according to the characteristics of each type of transcription factor binding site sequence; the convolutional neural network is used to extract DNA spatial structure features from the DNA spatial structure encoding matrix using the weight sharing characteristics of the convolution operation; The cube multi-layer perception layer includes sequence dimension multi-layer perceptron, modality dimension multi-layer perceptron and channel dimension multi-layer perceptron, which is used to capture the complex spatial interaction relationship of cross-modal features through the MLP mechanism of sequence, modality and channel dimensions, and integrate the multi-modal features of DNA sequence features and DNA spatial structure features; The output layer includes several multi-layer perceptrons, which are used to map the results of the fusion features into different categories of transcription factor binding sites, so as to realize the classification of different transcription factor binding site features.
10. A DeepTFBS model according to claim 9, characterized in that: The fully dynamic convolutional network consists of two dynamic convolution blocks, a downsampling layer and a linear layer; the dynamic convolution block includes a one-dimensional dynamic convolution, a RELU activation function, and a Dropout layer; the dynamic convolution includes a convergence-excitation multi-head attention module to achieve adaptive adjustment of the convolution kernel weights; The convolutional neural network includes two convolution blocks, a downsampling layer and a linear layer; the convolution block includes a one-dimensional convolution, a RELU activation function and a Dropout layer; the convolution has the characteristics of weight sharing and local receptive field, which is used to extract the features of DNA spatial structure with fewer parameters; The RELU activation function and Dropout layer enhance the generalization ability of the model by introducing nonlinear features.
Citation Information
Patent Citations
Transcription factor binding site prediction method based on depth convolution automatic encoder
CN111312329A
Deep learning-based transcription factor binding site positioning method
CN114758721A
TF-DNA combination recognition method based on multi-feature fusion
CN115810398A
Transcription factor binding site prediction method and system
CN118629495A
Method for predicting corn chromatin open region based on multi-scale convolutional network and gMLP
CN118737284A