A method for predicting the folding state of i-motifs based on deep learning

Automatically learns the complex structural characteristics of DNA sequences through the DeepIM model, solving the problems of high cost of traditional methods and insufficient accuracy of machine learning methods, and achieving efficient and interpretable i-motifs folding state prediction.

CN120048353BActive Publication Date: 2025-07-11YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510519006.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-11
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

Traditional methods are costly, time-consuming and difficult to apply at scale when predicting the folded state of i-motifs, and machine learning methods are difficult to capture complex feature interactions and insufficient prediction accuracy.

Method used

DeepIM model based on deep learning is adopted, combining position coding, channel attention and spatial attention mechanism, DNA sequences are processed through self-attention mechanism, complex structural features are automatically learned, and DeepIM model is constructed for prediction.

Benefits of technology

The prediction ability of i-motifs folded state and the interpretability of the model are significantly improved, the training efficiency and the ability to process long sequences are improved, and the performance and interpretability of the model are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048353B_ABST
    Figure CN120048353B_ABST
Patent Text Reader

Abstract

The present invention provides a method for predicting the folding state of i-motifs based on deep learning, which solves the problems that existing machine learning methods are difficult to capture complex feature interaction relationships and discover hidden laws, etc. The method includes the following steps: S1: Data acquisition and preprocessing; S2: Screening of i-motif candidate sequences; S3: Sequence encoding and feature extraction; S4: Construction and training of the DeepIM model; S5: Model evaluation; S6: Result visualization and comparison. The present invention has the advantages of good prediction effect, good interpretability, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of healthcare informatics, and particularly relates to a method for predicting the folding state of i-motifs based on deep learning. Background Art

[0002] i-motifs (iM) are four-stranded structures formed by cytosine-rich DNA sequences, and they play key roles in gene expression regulation, telomere maintenance, and the occurrence and development of various diseases. Especially in cancer research, the formation and stability of i-motifs are closely related to tumor suppressor genes. Traditional research methods for i-motifs rely on experimental techniques such as nuclear magnetic resonance (NMR) and X-ray crystallography. Although these methods are accurate, they are costly, time-consuming, and difficult to apply on a large scale. Machine learning methods including random forest, RUSBoost, XGBoost, naive Bayes, etc. can be used to predict the folding state of i-motifs, but they require manual selection and extraction of features, are difficult to capture complex feature interaction relationships and discover hidden rules, and also need to be improved in terms of prediction accuracy.

[0003] The present invention proposes a method for predicting the folding state of i-motifs based on deep learning, aiming to automatically learn complex structural features in the putative sequences by constructing a deep learning model, thereby improving the prediction ability of the folding state of i-motifs and enhancing the interpretability of the model. Summary of the Invention

[0004] The object of the present invention is to provide a method for predicting the folding state of i-motifs based on deep learning with reasonable design and good prediction effect in view of the above problems.

[0005] To achieve the above object, the present invention adopts the following technical solutions: A method for predicting the folding state of i-motifs based on deep learning, comprising the following steps:

[0006] S1: Data acquisition and preprocessing;

[0007] S2: Screening of i-motif candidate sequences;

[0008] S3: Sequence encoding and feature extraction;

[0009] S4: Construction and training of the DeepIM model;

[0010] S5: Model evaluation;

[0011] S6: Result visualization and comparison.

[0012] In the above method for predicting the folding state of i-motifs based on deep learning, step S1 includes the following steps:

[0013] S11: Data download, obtaining the i-motif forming sequence data of the HEK293T cell line from the NCBI GEO database, including three biological replicates;

[0014] S12: Format conversion and peak region definition, using SEACR v1.3 to convert BigWig to a bedGraph file, merging the peak regions of the three replicates, defining the overlapping i-motifs peak regions in the three biological replicates as the final high-confidence i-motifs peak regions, and defining the remaining parts of the overlapping regions as the spacer regions.

[0015] In the above method for predicting the folding state of i-motifs based on deep learning, in step S2, the Putative-iM-Searcher tool is used to search for potential i-motif sequences in the high-confidence peak regions and spacer regions respectively. The putative i-motifs searched in the high-confidence i-motifs peak regions are defined as folded i-motifs, and those searched in the spacer regions are defined as unfolded C-rich sequences.

[0016] In the above method for predicting the folding state of i-motifs based on deep learning, step S3 includes the following steps:

[0017] S31: Extract all 4-mer subsequences with a step size of 1;

[0018] S32: Add CLS and SEP tokens;

[0019] S33: Create an encoding dictionary including all possible 4-mers and special tokens;

[0020] S34: Convert each token to an index;

[0021] S35: Combine the indices into a numeric matrix.

[0022] In the above method for predicting the folding state of i-motifs based on deep learning, in step S4, the DeepIM model processes the input data through an embedding layer, a positional encoding layer, a local pyramid attention mechanism, and a Transformer encoder to generate the final prediction result.

[0023] In the above method for predicting the folding state of i-motifs based on deep learning, step S4 includes the following steps:

[0024] S41: The embedding layer converts the input numeric matrix into a high-dimensional vector;

[0025] S42: The positional encoding layer provides information about each position in the sequence for the model;

[0026] S43: The local pyramid attention mechanism combines the channel attention and spatial attention mechanisms. First, the channel attention mechanism is applied to generate channel attention weights and weight the input feature map X; then the spatial attention mechanism is applied to generate spatial attention weights and weight the feature map weighted by channel attention, and finally, the weighted feature map and channel attention weights are output;

[0027] S44: The Transformer encoder includes a multi-head self-attention mechanism and a feed-forward neural network. The multi-head self-attention mechanism generates a weighted feature representation, adds the input directly to the output of the multi-head self-attention, and performs layer normalization, and further processes the feature representation through the feed-forward neural network;

[0028] S45: The loss function uses weighted cross-entropy loss to calculate the difference between the model's prediction result and the true label;

[0029] S46: Backpropagation updates the model's parameters by calculating the gradient of the loss function.

[0030] In the above method for predicting the folding state of i-motifs based on deep learning, the calculation formula for positional encoding in step S42 is:

[0031] ;

[0032] ;

[0033] where, is the position vector, is the dimension, is the model dimension, and the output shape of the positional encoding is the same as the output shape of the embedding layer.

[0034] In the above method for predicting the folding state of i-motifs based on deep learning, the loss function formula in step S45 is:

[0035] ;

[0036] where, is the true label of the sample, is the predicted probability that the model belongs to the positive class for the sample, is the weight of the positive class, is the weight of the negative class.

[0037] In the above-mentioned method for predicting the folding state of i-motifs based on deep learning, step S5 includes calculating the accuracy rate and performance metrics, and the performance metrics include precision, recall, specificity, area under the ROC curve, and area under the PRC curve.

[0038] In the above-mentioned method for predicting the folding state of i-motifs based on deep learning, step S6 plots a curve graph for showing the precision-recall curve and the feature curve.

[0039] Compared with the existing technologies, the advantages of the present invention are as follows: The DeepIM model effectively captures the long-range dependence relationships in the sequence data through the self-attention mechanism. Applying this architecture to predict the folding state of the i-motif structure, the parallel processing ability of the DeepIM model greatly speeds up the training speed and improves the ability to process long sequences, significantly improving the training efficiency and model performance. Compared with the traditional machine learning models, the self-attention mechanism increases the interpretability of the model; Using the k-mer encoding method to encode the DNA i-motifs sequences simplifies the complexity of the data while retaining sufficient information for subsequent analysis. This encoding method not only reduces the data dimension, helps to improve the training efficiency of the model, but also enables the model to more effectively process and understand the DNA sequence data; Combining the channel attention and spatial attention mechanisms forms a new local pyramid attention mechanism LPA. This fusion not only improves the performance of the model but also provides a new direction for the research of the attention mechanism. LPA can simultaneously consider the importance of features and positions, providing more comprehensive sequence information for the model. Brief Description of the Drawings

[0040] Figure 1 is the workflow diagram of the present invention;

[0041] Figure 2 is the DeepIM architecture diagram of the present invention;

[0042] Figure 3 is the DeepIM five-fold cross-folding validation AUROC curve graph of the present invention;

[0043] Figure 4 is the comparison graph of the AUROC curves between DeepIM of the present invention and other machine learning models;

[0044] Figure 5 is the comparison graph of the precision-recall curves (PRC) between DeepIM of the present invention and other machine learning models;

[0045] Figure 6 is the visualization graph of the attention weights of the present invention. Detailed Embodiments

[0046] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0047] As Figure 1-6 shown, a method for predicting the folding state of i-motifs based on deep learning uses a deep learning model DeepIM based on the Transformer architecture to predict the folding state of DNA i-motifs structures. The model combines the Positional Encoding, Channel Attention, and Spatial Attention mechanisms to process the DNA i-motifs sequence data encoded by 4-mers. The specific steps are as follows:

[0048] S1: Data acquisition and preprocessing;

[0049] S2: Screening of i-motif candidate sequences;

[0050] S3: Sequence encoding and feature extraction;

[0051] S4: Construction and training of the DeepIM model;

[0052] S5: Model evaluation;

[0053] S6: Result visualization and comparison.

[0054] Specifically, step S1 includes the following steps:

[0055] S11: Data download, download the BigWig format data of the i-motifs forming sequences of the human embryonic kidney (HEK293T) cell line from the NCBI GEO database (accession number GSE220882), and this cell line has three biological replicate data;

[0056] S12: Format conversion and peak region definition, convert the downloaded BigWig file to a bedGraph file, and use SEACR v1.3 with parameters set to "0.01" and "non-stringent" to accumulate the i-motifs peak regions of the three files. Define the overlapping i-motifs peak regions in the three biological replicates as the final high-confidence i-motifs peak regions, and define the remaining parts after removing the overlapping regions as the spacer regions.

[0057] Specifically, in step S2, the Putative-iM-Searcher tool is used to search for potential i-motif sequences in the high-confidence peak region and the spacer region respectively. The putative i-motifs searched in the high-confidence i-motifs peak region are defined as folded i-motifs, and those searched in the spacer region are defined as unfolded C-rich sequences.

[0058] In addition, in step S3, all samples and labels are uniformly encoded using the k-mer encoding method. Set k to 4 and use the k-mer_encode function to convert the given i-motif sequence into a series of 4-mer tokens. For each input sequence, the function extracts all possible 4-mer subsequences from the sequence and adds the special tokens CLS_TOKEN and SEP_TOKEN to the beginning and end positions of the sequence. Then use the tokenize function to convert the generated 4-mers into their indices in the encoding dictionary, which maps each 4-mer to a unique integer. The final result is a digital matrix for subsequent prediction by the deep learning model, and the specific steps are as follows:

[0059] S31: Extract all 4-mer subsequences with a step size of 1;

[0060] S32: Add CLS and SEP tokens;

[0061] S33: Create an encoding dictionary including all possible 4-mers and special tokens;

[0062] S34: Convert each token to an index;

[0063] S35: Combine the indices into a digital matrix.

[0064] Meanwhile, in step S4, the DeepIM deep learning model is used to predict the encoded data. A model based on the Transformer architecture is used, which combines positional encoding, channel attention, and spatial attention mechanisms to predict the folding state of i-motifs in DNA sequences. The model processes the input data through an embedding layer, a positional encoding layer, an LPA layer, and a Transformer encoder to generate the final prediction result. During the model training and evaluation, a cross-entropy loss function, an Adam optimizer, and various performance metrics are used. Through these methods, the model can effectively predict the folding state of i-motifs, providing new tools and methods for i-motif prediction.

[0065] Visibly, step S4 includes the following steps:

[0066] S41: The embedding layer converts the input digital matrix into a high-dimensional vector. Each input 4-mer encoding is mapped into a vector space of a fixed dimension, which is called the model dimension (d_model). The output shape of the embedding layer is (sequence length, batch size, model dimension);

[0067] S42: The positional encoding layer is used to provide information about each position in the sequence for the model. The positional encoding is achieved by adding a position-related vector to the embedding vector;

[0068] S43: The Local Pyramid Attention mechanism (LPA) is a composite attention mechanism that combines Channel Attention and Spatial Attention, aiming to enhance the model's ability to perceive important features. This mechanism focuses on the channel dimension and spatial dimension of the features respectively, and can capture key information in the input data more comprehensively, thus improving the performance of the model;

[0069] S44: The Transformer encoder consists of multiple encoder layers. Each encoder layer includes a multi-head self-attention mechanism and a feed-forward neural network. The multi-head self-attention mechanism allows the model to learn information in parallel in different representation subspaces, while the feed-forward neural network is used to further process this information. The weighted feature representation is generated through the multi-head self-attention mechanism. The input is directly added to the output of the multi-head self-attention and layer normalization is performed. The feature representation is further processed through the feed-forward neural network. The dimension of the model, which is also the output dimension of the embedding layer, is set to 256, and the input and output dimensions are kept consistent to ensure the feasibility of the residual connection. The number of heads in the multi-head self-attention mechanism is set to 8, enabling the model to capture more details; Residual connection and layer normalization: The input is directly added to the output of the feed-forward neural network and layer normalization is performed. The number of encoder layers is set to 3, and the number of layers determines the depth of the Transformer encoder. The more layers there are, the richer the context information the model can capture, but the computational complexity and training time will also increase. The dimension of the middle layer of the feed-forward neural network is set to 512 to enhance the non-linear expression ability of the model. To prevent the model from overfitting, Dropout is added, with the parameter set to 0.5, and a part of the neurons are randomly discarded during the training process for regularization; Finally, a residual connection is used and layer normalization is performed to alleviate the vanishing gradient problem in the deep network and accelerate the training process. After being processed by all encoder layers, the output layer converts the output of the Transformer encoder into the final prediction result through a linear layer;

[0070] S45: The loss function uses weighted cross-entropy loss (CrossEntropyLoss) to calculate the difference between the model's prediction results and the true labels. Weighted cross-entropy loss is an effective method for dealing with class imbalance problems in binary classification. By assigning different weights to samples of different classes, the contributions of different classes can be balanced, thereby improving the model's prediction performance for the minority class. The output of the loss function is a scalar value representing the prediction error of the model;

[0071] S46: Backpropagation updates the model's parameters by calculating the gradient of the loss function. The optimizer uses the Adam optimizer to optimize the model's training process by adjusting the learning rate and weight decay parameters. The optimizer updates the model's parameters at each training step to minimize the value of the loss function. The optimization process includes adjusting hyperparameters such as the learning rate, weight decay, and batch size to improve the training effect of the model.

[0072] Obviously, the calculation formula for positional encoding in step S42 is:

[0073] ;

[0074] ;

[0075] Among them, is the position vector, is the dimension, is the model dimension, and the output shape of the positional encoding is the same as the output shape of the embedding layer.

[0076] Specifically, the core idea of the channel attention mechanism in step S43 is to extract the global information of the input feature map through global pooling operations (such as global average pooling and global max pooling), and then generate channel attention weights through two fully connected layers (FC) and a sigmoid activation function. These weights are used to weight each channel of the input feature map to enhance the feature representation of important channels. Global average pooling and global max pooling are performed on the input feature map X to obtain two one-dimensional feature vectors avg_pool and max_pool respectively; avg_pool and max_pool are respectively passed through two fully connected layers (FC1 and FC2), where the output dimension of FC1 is 1 / 8 of the input dimension, and the output dimension of FC2 is the same as the input dimension. The output obtained is passed through the sigmoid activation function to generate the channel attention weights. Finally, the generated channel attention weights are multiplied by the input feature map X to obtain the weighted feature map.

[0077] The core idea of the spatial attention mechanism is to reduce the number of channels of the input feature map to a fixed value through a 1x1 convolutional layer, and then extract spatial information through global average pooling and global max pooling to generate spatial attention weights. These weights are used to weight each position of the input feature map to enhance the feature representation of important positions. First, perform 1x1 convolution on the input feature map X to reduce the number of channels to 256, and then perform global average pooling and global max pooling on the convolved feature map to obtain two one-dimensional feature vectors avg_pool and max_pool respectively. Concatenate avg_pool and max_pool into a feature vector. Increase the number of channels of the concatenated feature vector to the same dimension as the input feature map for dimension matching, and the output generates spatial attention weights through the sigmoid activation function. Multiply the generated spatial attention weights by the input feature map X to obtain the weighted feature map.

[0078] The LPA layer combines the channel attention and spatial attention mechanisms. First, apply the channel attention mechanism to generate channel attention weights and weight the input feature map X; then apply the spatial attention mechanism to generate spatial attention weights and weight the feature map weighted by the channel attention, and finally output the weighted feature map and the channel attention weights.

[0079] Preferably, the loss function formula in step S45 is:

[0080] ;

[0081] where is the true label of the sample, is the predicted probability that the model belongs to the positive class, is the weight of the positive class, is the weight of the negative class.

[0082] Obviously, step S5 evaluates the model, and the steps include calculating the accuracy and other performance metrics. The accuracy calculates the matching degree between the model prediction result and the true label, and generates a percentage value indicating the prediction accuracy of the model. The performance metrics include Precision, Recall, Specificity, the area under the ROC curve (ROCAUC), and the area under the PRC curve (PRCAUC). These metrics are used to evaluate the performance of the model, especially the effect when dealing with imbalanced data.

[0083] In addition, step S6 plots a curve graph to display the Precision-Recall Curve (PRC) and the Receiver Operating Characteristic Curve (ROC). The PRC curve shows the changes in precision and recall of the model at different thresholds, and the ROC curve shows the changes in the true positive rate and false positive rate of the model at different thresholds. The Area Under the Curve (AUC) represents the performance of the model, and the larger the AUC value, the better the performance of the model.

[0084] As Figure 3 shown, the five-fold cross-validation method can comprehensively evaluate the classification performance of the model, improve the reliability of model evaluation, enhance model interpretability, and the average AUROC = 0.94 proves that the model has achieved a good classification effect.

[0085] As Figure 4 shown, the AUROC value obtained by DeepIM on the test set reaches 0.96, which is higher than that of the Naive Bayes model, Linear Discriminant Analysis model, Easy Ensemble model, RUSBoost model, XGBoost model, BalancedRandomForest model, etc., and has the best prediction effect.

[0086] As Figure 5 shown, the AUPRC value obtained by DeepIM on the test set reaches 0.80, which is higher than that of the Naive Bayes model, Linear Discriminant Analysis model, Easy Ensemble model, RUSBoost model, XGBoost model, BalancedRandomForest model, etc., and has the best prediction effect.

[0087] IM-Seeker is a machine learning-based computational tool that first identifies potential i-motif-forming sequences from the input DNA sequences through Putative-iM-Searcher, then manually extracts 33 features from the sequences to construct a classification model, and combines the BalancedRandomForest classifier to predict the folding status of i-motifs. Compared with iM-Seeker, the DeepIM model can perform end-to-end learning without manually extracting features. The model can automatically extract useful features from the original sequences, reducing the complexity of feature engineering. Moreover, LPA can capture local and global sequence features, and Transformer is good at dealing with long-range dependencies, which gives them potential advantages in dealing with complex DNA sequence structures. In addition, the prediction process of the model can be explained through attention weights, enhancing the interpretability of the model. The following table shows the comparison results of the prediction performance between DeepIM and IM-seeker. It can be seen that DeepIM has significantly improved in terms of accuracy, AUC, recall rate, and specificity, demonstrating the advantages of the DeepIM model in predicting the folding status of i-motifs.

[0088] model accuracy AUC recall rate specificity DeepIM 0.92 0.95 0.86 0.88 im-seeker 0.90 0.87 0.83 0.87

[0089] DeepIM provides the interpretability of the model through the attention mechanism. The multi-head self-attention mechanism of the Transformer model allows the model to learn information in parallel in different representation subspaces and capture the relationships between different positions in the input sequence. By visualizing the weights of each attention head, it can be understood how the model pays attention to different positions in the input sequence. The LPA module combines ChannelAttention and SpatialAttention, and extracts the global and local information of the input features through convolution and global pooling operations, which is used to enhance the model's perception ability of important features in the input sequence. By visualizing the attention weights, it can be understood which feature dimensions of the input sequence the model pays more attention to and which nucleotides or k-mers in the input sequence the model pays more attention to.

[0090] As Figure 6 shown, it can be seen that the model assigns higher weights to multiple consecutive C-loops, indicating that it can identify these key regions and use them as important features for predicting the folding status of i-motifs. Since i-motifs rely on consecutive C-tracts to stabilize the structure through semi-protonated C:C base pairs (C+:C), it can be verified that the prediction of the model is consistent with known biological knowledge, indicating that the prediction of the model is reasonable.

[0091] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art to which the present invention pertains can make various modifications or supplements to the described specific embodiments or use similar means for substitution, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

[0092] Although terms such as deep learning and DeepIM model are used more frequently herein, the possibility of using other terms is not excluded. The use of these terms is only for more convenient description and explanation of the essence of the present invention; interpreting them as any additional limitation is contrary to the spirit of the present invention.

Claims

1. A method for predicting the folding state of i-motifs based on deep learning, characterized in that, It includes the following steps: S1: Data acquisition and preprocessing; S2: Screening of i-motif candidate sequences; S3: Sequence encoding and feature extraction; S31: Extract all 4-mer subsequences with a step size of 1; S32: Add CLS and SEP markers; S33: Create an encoding dictionary including all possible 4-mers and special markers; S34: Convert each token to an index; S35: Combine the indices into a numeric matrix; S4: Construction and training of the DeepIM model. The DeepIM model processes the input data through an embedding layer, a positional encoding layer, a local pyramid attention mechanism, and a Transformer encoder to generate the final prediction result; S41: The embedding layer converts the input numeric matrix into a high-dimensional vector; S42: The positional encoding layer provides information about each position in the sequence for the model; S43: The local pyramid attention mechanism combines the channel attention and spatial attention mechanisms. First, it applies the channel attention mechanism to generate channel attention weights and weight the input feature map X; then it applies the spatial attention mechanism to generate spatial attention weights and weight the feature map weighted by the channel attention, and finally outputs the weighted feature map and the channel attention weights; S44: The Transformer encoder includes a multi-head self-attention mechanism and a feed-forward neural network. The multi-head self-attention mechanism generates a weighted feature representation, adds the input directly to the output of the multi-head self-attention, and performs layer normalization, and further processes the feature representation through the feed-forward neural network; S45: The loss function uses weighted cross-entropy loss to calculate the difference between the model prediction result and the true label; S46: Backpropagation updates the model parameters by calculating the gradient of the loss function; S5: Model evaluation; S6: Result visualization and comparison.

2. The folding state prediction method of i-motifs based on deep learning according to claim 1, wherein The step S1 includes the following steps: S11: Data download. Obtain the i-motif forming sequence data of the HEK 293T cell line from the NCBI GEO database, including three biological replicates; S12: Format conversion and peak region definition. Use SEACR v1.3 to convert BigWig to a bedGraph file, merge the peak regions of the three replicates, define the overlapping i-motifs peak regions in the three biological replicates as the final high-confidence i-motifs peak regions, and define the remaining parts of the overlapping regions as the spacer regions.

3. A method for predicting the folding state of i-motifs based on deep learning according to claim 1, characterized in that, In the step S2, the Putative-iM-Searcher tool is used to search for potential i-motif sequences in the high-confidence peak regions and spacer regions respectively. The putative i-motifs searched in the high-confidence i-motifs peak regions are defined as folded i-motifs, and those searched in the spacer regions are defined as unfolded C-rich sequences.

4. A method for predicting the folding state of i-motifs based on deep learning according to claim 1, characterized in that, The calculation formula for positional encoding in the step S42 is: ; ; Among them, is the position vector, is the dimension, is the model dimension, and the output shape of the positional encoding is the same as the output shape of the embedding layer.

5. A method for predicting the folding state of i-motifs based on deep learning according to claim 1, characterized in that, The loss function formula in the step S45 is: ; Among them, is the true label of the sample, is the predicted probability that the model assigns the sample to the positive class, is the weight of the positive class, is the weight of the negative class.

6. The folding state prediction method of i-motifs based on deep learning according to claim 1, characterized in that, The step S5 includes calculating the accuracy and performance metrics. The performance metrics include precision, recall, specificity, area under the ROC curve, and area under the PRC curve.

7. A method for predicting the folding state of i-motifs based on deep learning according to claim 6, characterized in that, The described step S6 of drawing a curve graph is used to display the precision-recall curve and the characteristic curve.

Citation Information

Patent Citations

  • DNA synthesis difficulty prediction system and application thereof

    CN116312783A

  • Multi-modal attention deep learning method for enhancing virus recognition in metagenome data

    CN118918954A