A RNA G-quadruplex prediction method and system based on pre-trained model and RNA secondary structure

By using a method based on pre-trained models and RNA secondary structure in RNA G-quadrilateral prediction, positive and negative samples are constructed and sequence characteristics and secondary structure characteristics are extracted, the problems of identifying non-canonical rG4 and not fully considering RNA secondary structure in the prior art are solved, and more efficient and accurate rG4 prediction is achieved.

CN119724349BActive Publication Date: 2025-05-16YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510228817.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-16
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The existing RNA G-quadruple (rG4) prediction methods lack flexibility, difficult to identify non-canonical or variant forms of rG4, and do not fully consider the secondary structural characteristics of RNA, resulting in insufficient comprehensive prediction.

Method used

Using a method based on pre-trained models and RNA secondary structure, we use positive and negative samples to construct sequence characteristics and secondary structural characteristics of RNA sequences, and combine BERT models and neural network models for model training to improve the prediction performance of rG4.

Benefits of technology

It significantly improves the accuracy and performance of rG4 prediction, can better capture complex rG4 sequence characteristics and secondary structural characteristics of RNA, and provide more comprehensive and accurate prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119724349B_ABST
    Figure CN119724349B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for predicting RNA G-quadruplex based on a pre-training model and RNA secondary structure, including obtaining the position information of human rG4 on human transcripts; for each sequence, padding the same length to both sides according to its sequence position coordinates so that the total length reaches a set length value; obtaining human cDNA sequence data as a reference sequence, extracting rG4 data containing flanking sequence information from the cDNA sequence according to the filled sequence coordinates as a positive sample sequence; shuffling each positive sample sequence to obtain a negative sample sequence; generating RNA secondary structure features for each sample sequence; extracting sequence features of the sample sequence using a pre-training model; and inputting sequence features and RNA secondary structure features into a prediction model for model training. This scheme utilizes the secondary structure features of RNA sequences and introduces the secondary structure features as auxiliary information, which can significantly improve the prediction performance of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of computer bioinformatics, and in particular relates to an RNA G-quadruplex prediction method and system based on a pre-training model and RNA secondary structure. Background Art

[0002] Existing RNA G-quadruplex (rG4) prediction methods include regular expression matching-based methods (QGRSmapper), scoring-based methods (cGcC scoring, G4Hunter), and neural network-based methods (G4NN, rG4detector).

[0003] The method based on regular expression matching lacks flexibility and can only identify canonical rG4 motifs that conform to specific patterns, and has limited recognition ability for non-canonical or variant forms of rG4. Similarly, methods that rely on scoring schemes still focus on this canonical rG4 motif, ignoring more complex and variable sequence features, and are prone to false positive results because they tend to mark all G-rich regions as potential rG4 formation regions. In contrast, G4NN predicts the formation of rG4 through a neural network model, which can theoretically better capture the complex rG4 sequence features. However, due to its small training data set (less than 500 verified rG4 reference sequences), its accuracy in practical applications is still limited. In addition, the research focus of rG4detector as a deep learning method is on predicting the stability of rG4 structure, so its performance in rG4 classification tasks is relatively limited. It is worth noting that RNA, as a single-stranded molecule, has a strong folding ability and can form complex secondary structures. However, none of the methods disclosed in the prior art fully consider this important characteristic, resulting in an incomplete prediction of rG4 formation. Summary of the invention

[0004] The purpose of the present invention is to propose a method for predicting RNA G-quadruplex based on a pre-training model and RNA secondary structure in response to the problems existing in the prior art;

[0005] Another object of the present invention is to propose an RNA G-quadruplex prediction system based on a pre-training model and RNA secondary structure to address the problems existing in the prior art.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] A method for predicting RNA G-quadruplex based on a pre-trained model and RNA secondary structure, the method comprising:

[0008] Construct positive and negative samples:

[0009] Obtain the location information of human rG4 on human transcripts;

[0010] For each sequence, pad the two sides with the same length according to its sequence position coordinates so that the total length reaches the set length value;

[0011] Obtain human cDNA sequence data as a reference sequence, and extract rG4 data containing flanking sequence information from the cDNA sequence according to the filled sequence coordinates as a positive sample sequence;

[0012] The negative sample sequence is obtained by shuffling each positive sample sequence while keeping the dinucleotide frequency unchanged;

[0013] Generate RNA secondary structure features for each positive sample sequence and negative sample sequence;

[0014] Use the pre-trained model to extract sequence features of positive sample sequences and negative sample sequences;

[0015] The sequence features and RNA secondary structure features are input into the prediction model for model training.

[0016] In the above-mentioned RNA G-quadruplex prediction method based on the pre-training model and RNA secondary structure, the set length value is 125 nt.

[0017] In the above-mentioned RNA G-quadruplex prediction method based on the pre-training model and RNA secondary structure, the same length is padded on both sides according to the sequence position coordinates so that the total length reaches the set length value. If the padded sequence coordinates exceed the boundary of the corresponding transcript, the excess part is padded with arbitrary nucleotides;

[0018] After extracting rG4 data containing flanking sequence information from the cDNA sequence, the rG4 data is deduplicated to obtain the desired positive sample.

[0019] In the above-mentioned RNA G-quadruplex prediction method based on pre-trained model and RNA secondary structure, the pre-trained BERT model is used to extract sequence features of positive sample sequences and negative sample sequences, specifically including:

[0020] Convert positive sample sequences and negative sample sequences into tokens required by BERT;

[0021] BERT outputs sequence features for each sample sequence based on the input tokens.

[0022] In the above-mentioned RNA G-quadruplex prediction method based on the pre-training model and RNA secondary structure, the sequence feature is a high-dimensional vector, and the secondary structure feature is a low-dimensional vector;

[0023] The prediction model reduces the dimensionality of sequence features and increases the dimensionality of secondary structure features so that the sequence features and secondary structure features have the same dimension.

[0024] In the above-mentioned RNA G-quadruplex prediction method based on a pre-trained model and RNA secondary structure, the prediction model includes a first convolutional network, a hierarchical multi-scale residual network, a second convolutional neural network, layer normalization, a self-attention mechanism, a multi-layer perceptron, a pooling layer and an activation function;

[0025] The second convolutional neural network includes an SE module, which compresses each feature into a scalar value. The output of the SE module is element-by-element multiplied with the original feature to recalibrate the feature. Finally, the recalibrated feature is input into the normalization layer, self-attention mechanism, multi-layer perceptron, pooling layer and activation function for further processing.

[0026] In the above-mentioned RNA G-quadruplex prediction method based on pre-trained model and RNA secondary structure, the first convolutional neural network is used to reduce and increase the dimension of sequence features and secondary structure features to 128 dimensions respectively.

[0027] In the above-mentioned RNA G-quadruplex prediction method based on the pre-trained model and RNA secondary structure, the prediction model uses cross entropy loss as the loss function of the model for model training;

[0028] The RNA secondary structure characteristics of each positive sample sequence and negative sample sequence include the unpaired probability value of each nucleotide in the corresponding sequence.

[0029] A method for predicting RNA G-quadruplexes based on a pre-trained model and RNA secondary structure, wherein the prediction model obtained by the above method is executed to predict RNA G-quadruplexes:

[0030] Outputting a prediction of whether the RNA sequence to be predicted contains rG4 based on the input RNA sequence data to be predicted;

[0031] The RNA sequence data includes the sequence characteristics and RNA secondary structure characteristics of the RNA sequence to be predicted;

[0032] The sequence features are obtained by encoding the RNA sequence to be predicted by a pre-trained model;

[0033] The RNA secondary structure feature is the unpairing probability value of each nucleotide in the RNA sequence to be predicted.

[0034] An RNA G-quadruplex prediction system based on a pre-trained model and RNA secondary structure, comprising a prediction module and a feature extraction module;

[0035] The prediction module is used to output a prediction of whether the RNA sequence to be predicted contains rG4 based on the input RNA sequence data to be predicted;

[0036] The RNA sequence data includes sequence features and RNA secondary structure features of the RNA sequence to be predicted;

[0037] The feature extraction module includes a pre-training module and an RNA secondary structure acquisition module. The pre-training module is used to encode the RNA sequence to be predicted to obtain the sequence feature, and the RNA secondary structure acquisition module is used to extract the unpairing probability value of each nucleotide in the RNA sequence to be predicted to obtain the RNA secondary structure feature.

[0038] The advantages of the present invention are:

[0039] 1) RNA, as a single-stranded molecule, has strong folding ability and can form complex secondary structures. However, the currently disclosed technologies have not fully considered this important characteristic. This scheme uses the secondary structure characteristics of RNA sequences and introduces secondary structure characteristics as auxiliary information. Compared with the existing models, it can significantly improve the prediction performance of the model;

[0040] 2) This scheme uses the known position information of rG4 on the human transcript to determine the sequence coordinates of each positive sequence by flanking the position, and then extracts the sequence corresponding to the sequence coordinates from the cDNA sequence as a positive sample. At the same time, the BERT model is used to better capture the characteristics of contextual features. Combined with the use of a pre-trained BERT model to encode the sequence, the influence of the flanking sequence of rG4 on the formation of rG4 is taken into account, thereby significantly improving the prediction performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 Shown is a flow chart of a method for predicting RNA G-quadruplex based on a pre-trained model and RNA secondary structure according to an embodiment of the present invention;

[0042] Figure 2 FIG. 1 is a model framework diagram of an RNA G-quadruplex prediction method based on a pre-trained model and RNA secondary structure according to an embodiment of the present invention;

[0043] Figure 3 The figure shows a framework diagram of a neural network model constructed in the RNA G-quadruplex prediction method based on a pre-training model and RNA secondary structure according to an embodiment of the present invention;

[0044] Figure 4 It is a flow chart showing the RNA G-quadruplex prediction system based on the pre-trained model and RNA secondary structure in an embodiment of the present invention;

[0045] Figure 5 Shown is a comparison of the ROC curves of each method;

[0046] Figure 6 This is a t-SNE visualization diagram of several core modules in the neural network model constructed in the RNA G-quadruplex prediction method based on the pre-training model and RNA secondary structure in an embodiment of the present invention. DETAILED DESCRIPTION

[0047] This protocol provides a method for RNA G-quadruplex prediction based on a pre-trained model and RNA secondary structure. Figure 1 and Figure 2 As shown, the method includes:

[0048] The first step is to construct a positive sample data set: human rG4 data from rG4-seq (K+) were downloaded from the G4Atlas database. The data contains the position information of rG4 on human transcripts. For each sequence, the same length was padded on both sides according to its sequence position coordinates to make its total length 125nt. Then, human cDNA sequence data was downloaded from the Ensembl database as a reference sequence, and rG4 data containing flanking sequence information was extracted from the cDNA sequence according to the padded sequence coordinates.

[0049] If the padded sequence coordinates exceed the boundaries of the corresponding transcript, the excess part is filled with N (representing any nucleotide).

[0050] Finally, the CD-HIT tool is used to deduplicate the processed sequences to obtain the final positive samples.

[0051] Experimentally verified rG4 data were obtained from the G4Atlas database to ensure data reliability and accuracy. Using the known location information of rG4 on human transcripts, the sequence coordinates of each positive sequence were determined by filling the flanks at the position, and then the sequence corresponding to the sequence coordinates was extracted from the cDNA sequence as a positive sample. By filling the flank sequences, more contextual information was provided, and the effect of the flanking sequence of rG4 on the formation of rG4 was taken into account, which helped the model better understand the formation mechanism of rG4.

[0052] The second step is to construct a negative sample data set: for each positive sample, use the dishuffle tool (to shuffle the RNA sequence data while keeping the dinucleotide frequency unchanged) to process it to obtain the corresponding negative sample sequence. The final number of negative samples is the same as that of positive samples. These negative samples are consistent with the positive samples in dinucleotide frequency, but do not contain rG4 structure.

[0053] Step 3: Generate the RNA secondary structure information of positive and negative samples: Use the RNAplfold package in the ViennaRNA toolkit to generate the unpaired probability values of each nucleotide in each sequence through the command "RNAplfold -u1 <path / to / data>", which is used to reflect the RNA secondary structure information of the sequence.

[0054] The unpaired probability value provides the probability that each nucleotide in the RNA sequence does not participate in pairing, which helps the model understand the secondary structure characteristics of RNA. And this information can be used as auxiliary features to enhance the model's prediction ability for rG4 formation.

[0055] Step 4: Divide the dataset: Sort the sequences (in FASTA format) according to their description line information, and select the first 2050 (about 15%) sequences with a positive-negative sample ratio of 1:1 as the independent test set for model evaluation. The remaining sequences are randomly shuffled and divided into a training set and a validation set according to an 8:2 ratio for model training and tuning respectively.

[0056] Step 5: Feature extraction: Use a pre-trained BERT model to encode the sequences. First, divide the sequences into tokens in the form of k-mer (k = 3), and at the same time add special start tokens ([CLS]) at the beginning of these tokens and end tokens ([SEP]) at the end to represent the overall features and delimiters of the sequences, and also regard them as tokens. The pre-trained BERT model used is the already trained BERT model. In this solution, this BERT model is used for preliminary feature extraction, and the BERT model does not participate in the training during the training process of the prediction model.

[0057] Then, use the tokenizer to convert these tokens into the input format required by the BERT model (input tensors including tokenids and attention mask, etc.). Finally, input these encoded tokens into the pre-trained BERT model for processing to generate the context representation BERT embeddings of each token, which are the sequence features of the dataset. Among them, in this embodiment, each token is represented as a 768-dimensional high-dimensional vector.

[0058] The positive samples obtained by the first step of processing include the rG4 core region and its flanking sequences. These flanking sequences provide rich sequence background information for the rG4 in the core region. Through the context representation generated by the BERT model, these background information can be converted into high-dimensional feature vectors, so as to better understand the role of flanking sequences. For example, some flanking sequences may promote or inhibit the formation of rG4. The BERT model can capture these subtle differences and describe the characteristics of the sequence more comprehensively. At the same time, these feature vectors are fused with secondary structure features (such as unpaired probability values) to further enhance the model's feature representation capabilities, thereby making more accurate predictions.

[0059] Step 6: Building a neural network model: Figure 2 As shown in the figure, first, the first convolutional neural network (CNN) is used to increase and reduce the dimension of the sequence features (768 dimensions) and the secondary structure features (1 dimension) to 128 dimensions respectively to obtain the sequence feature graph and the structure feature graph. Secondly, the hierarchical multi-scale residual network is used to further extract the sequence features and the secondary structure features, and the extracted features are spliced.

[0060] Next, the concatenated features are input into the second convolutional neural network (CNN) for feature fusion, and the Squeeze-and-Excitation (SE) module is introduced to enhance the performance of CNN. The core idea of ​​the SE module is to enable the network to adaptively learn the importance of each channel and reweight the channel features according to these importance, thereby improving the representation ability of the model.

[0061] Specifically, the SE module first compresses the feature map of each channel into a scalar value through a global average pooling operation. This process can be seen as aggregating the global information of the entire feature map to capture the response of each channel on the entire image. Subsequently, the SE module builds a small-scale sub-network through two fully connected layers and an activation function to learn the dependencies between channels. The sub-network outputs a set of weights that represent the importance of each channel. Finally, the SE module recalibrates the features by performing element-by-element multiplication of these weights with the original feature map.

[0062] The feature map processed by the SE module can highlight the channels that are critical to the current task while suppressing unimportant channels. This dynamic adjustment mechanism enables the model to better focus on key features, thereby significantly improving its representation ability and final performance.

[0063] Finally, the output of the SE module is processed by layer normalization, self-attention mechanism and multi-layer perceptron (MLP) in turn, and the input is added to the processed features through residual connection. Residual connection ensures the effective transmission of information and avoids the gradient vanishing problem in deep networks. Subsequently, the features after residual connection are globally average pooled in the feature dimension to aggregate feature information. Global average pooling captures the global response of the entire sequence by averaging the feature maps of each channel while reducing the dimension of the features. The pooled result is flattened into a one-dimensional vector and passed to a fully connected layer with 1 output neuron.

[0064] After the fully connected layer, the sigmoid activation function is applied for processing, and its formula is as follows:

[0065] (1)

[0066] Through the sigmoid function, the model outputs a probability value between 0 and 1, representing the probability of containing rG4 in a sequence.

[0067] Step 7: Model training, tuning and evaluation: Use the divided training set to train the model and initialize the parameters using the Kaiming initialization method. Use AdamW as the optimizer of the model and use the cross entropy loss as the loss function of the model. The formula is as follows:

[0068] (2)

[0069] in, is the true label of the i-th sequence sample, is the model's predicted value for the i-th sequence sample, and N is the total number of samples in the batch.

[0070] After constructing and training the prediction model in the above manner, the prediction model is deployed to obtain an RNA G-quadruplex prediction system based on the pre-trained model and RNA secondary structure. The system includes a prediction module (for implementing the prediction model constructed by the above method) and a feature extraction module. The prediction module is used to output a prediction on whether the input RNA sequence to be predicted contains rG4; the RNA sequence data includes the sequence features and RNA secondary structure features of the RNA sequence to be predicted; the feature extraction module includes a pre-training module (for implementing the BERT model mentioned above) and an RNA secondary structure acquisition module. The pre-training module is used to encode the RNA sequence to be predicted to obtain sequence features, and the RNA secondary structure acquisition module is used to extract the unpaired probability value of each nucleotide in the RNA sequence to be predicted to obtain RNA secondary structure features. As Figure 4 shown, after constructing the prediction model, the specific prediction method is as follows:

[0071] Receive the RNA sequence to be predicted;

[0072] The pre-trained model encodes the RNA sequence to be predicted to obtain sequence features;

[0073] Use the RNAplfold package to generate the unpaired probability value of each nucleotide in the RNA sequence to be predicted through the "RNAplfold -u1 <path / to / data" command to obtain RNA secondary structure features;

[0074] Preferably, when the RNA sequence to be predicted received is not 125 nt, if the length of the RNA sequence to be predicted is less than 125 nt, it is filled to 125 nt length as a new RNA sequence to be predicted by taking the input RNA sequence to be predicted as the center according to the transcript for subsequent prediction. If the length of the RNA sequence to be predicted is greater than 125 nt, the sequence is trimmed to 125 nt length by retaining the middle part sequence as a new RNA sequence to be predicted and input to subsequent prediction.

[0075] The RNA secondary structure features and sequence features are fused into the RNA sequence data to be predicted and input into the prediction model;

[0076] The prediction model outputs a prediction on whether the RNA sequence to be predicted contains rG4 based on the input RNA sequence data to be predicted.

[0077] Use the validation set to monitor the training of the model, adjust the hyperparameters according to the performance of the model on the training set, and select the model with the best performance on the validation set. At the same time, use the early stopping method to prevent the model from overfitting. When the loss on the validation set does not decrease within 5 consecutive training cycles (epochs), stop the model training and save this model as the best model for evaluation on the independent test set.

[0078] In order to verify the performance of the method implemented in this scheme, the performance of the method implemented in this scheme is compared with existing methods, such as G4Hunter, cGcC scoring, G4NN, and rG4detector. The area under the ROC curve (AUROC) is mainly used as the performance evaluation indicator of the model, and the ROC curve is drawn. The results are shown in the figure. Figure 5 As shown, we can see that the method of this scheme has the highest AUC value, which proves that the classification performance of this scheme method is the best.

[0079] t-SNE visualization of core modules. I made t-SNE visualization of the processing results of several core modules of our method, and the results are as follows Figure 6 As shown in the figure, from the t-SNE visualization, we can see that the core modules of this solution have played their roles respectively. After each module, the model's ability to distinguish between positive and negative samples has been significantly improved.

[0080] The specific embodiments described herein are merely examples of the spirit of the present invention. Those skilled in the art of the present invention may make various modifications or additions to the specific embodiments described or replace them in a similar manner, such as changing the database to other feasible databases, replacing the above-mentioned tools with other tools having the same functions, etc., but they will not deviate from the spirit of the present invention or exceed the scope defined by the attached claims.

[0081] Although the terms human transcript, human rG4, human cDNA, positive sample, negative sample, RNA secondary structure feature, prediction model, pre-training model, etc. are used more frequently in this article, the possibility of using other terms is not excluded. The use of these terms is only to more conveniently describe and explain the essence of the present invention; interpreting them as any additional limitation is contrary to the spirit of the present invention.

Claims

1. A method for predicting RNA G-quadruplexes based on a pre-trained model and RNA secondary structure, characterized in that: The method includes: Constructing positive and negative samples: Obtaining the position information of human rG4 on human transcripts; For each sequence, pad the sequence with the same length on both sides according to its sequence position coordinates until the total length reaches a set length value; Obtaining human cDNA sequence data as a reference sequence, and extracting rG4 data containing flanking sequence information from the cDNA sequence according to the padded sequence coordinates as positive sample sequences; Obtaining negative sample sequences by shuffling each positive sample sequence while keeping the dinucleotide frequency unchanged; Generating the RNA secondary structure features of each positive sample sequence and negative sample sequence, and the secondary structure features include the unpaired probability of each nucleotide in the corresponding sequence, which are generated by using the RNAplfold package in the viennaRNA toolkit through the command "RNAplfold -u1 <path / to / data>"; Using a pre-trained BERT model to extract the sequence features of positive sample sequences and negative sample sequences, specifically including: Converting positive sample sequences and negative sample sequences into tokens required by BERT: Dividing the sequences into k-mer form tokens, adding special start markers at the beginning of the tokens, and adding end markers at the end to represent the overall features and delimiters of the sequences, and also treating them as tokens; Using a tokenizer to convert the tokens into the input format required by the BERT model; The pre-trained BERT generates the context representation BERT embeddings of each token based on the input tokens to obtain the sequence features of the dataset; Inputting the sequence features and RNA secondary structure features into a prediction model for model training, including further feature extraction of the sequence features and RNA secondary structure features respectively using a hierarchical multi-scale residual network, splicing the extracted features, inputting the spliced features into a second convolutional neural network for feature fusion, realizing the recalibration of the features through an SE module, aggregating the feature information through residual connection and global average pooling, and finally predicting RNA G-quadruplex based on the feature information.

2. The RNA G-quadruplex prediction method based on a pre-training model and RNA secondary structure according to claim 1, characterized in that: The set length value is 125nt.

3. The RNA G-quadruplex prediction method based on a pre-training model and RNA secondary structure according to claim 1, characterized in that: When padding the sequence with the same length on both sides according to its sequence position coordinates until the total length reaches the set length value, if the padded sequence coordinates exceed the boundary of the corresponding transcript, fill the exceeded part with any nucleotide; After extracting rG4 data containing flanking sequence information from the cDNA sequence, deduplicate the rG4 data to obtain the required positive samples.

4. The RNA G-quadruplex prediction method based on a pre-trained model and RNA secondary structure according to claim 1, characterized in that: The sequence features are high-dimensional vectors, and the secondary structure features are low-dimensional vectors; The prediction model reduces the dimension of the sequence features and increases the dimension of the secondary structure features so that the sequence features and secondary structure features have the same dimension.

5. The RNA G-quadruplex prediction method based on a pre-trained model and RNA secondary structure according to claim 1, characterized in that: The prediction model includes the hierarchical multi-scale residual network, the second convolutional neural network, and the first convolutional neural network; The second convolutional neural network includes the SE module, which compresses each feature into a scalar value. The output of the SE module is element-by-element multiplied with the original feature to recalibrate the feature. Finally, the recalibrated feature is input into the normalization layer, the self-attention mechanism, and the multi-layer perceptron. The result of the pooling of feature information is aggregated through residual connection and global average pooling and is passed to the fully connected layer and the sigmoid activation function for processing to predict RNA G-quadruplex.

6. The RNA G-quadruplex prediction method based on a pre-trained model and RNA secondary structure according to claim 5, characterized in that: The first convolutional neural network is used to reduce and increase the dimension of sequence features and secondary structure features to 128 dimensions respectively.

7. The RNA G-quadruplex prediction method based on a pre-training model and RNA secondary structure according to any one of claims 1 to 6, characterized in that: The prediction model uses cross entropy loss as the loss function of the model for model training.

8. A method for predicting RNA G-quadruplexes based on a pre-trained model and RNA secondary structure, characterized in that: RNA G-quadruplex prediction is performed by executing the prediction model obtained by the method according to any one of claims 1 to 7: Outputting a prediction of whether the RNA sequence to be predicted contains rG4 based on the input RNA sequence data to be predicted; The RNA sequence data includes the sequence characteristics and RNA secondary structure characteristics of the RNA sequence to be predicted; The sequence features are obtained by encoding the RNA sequence to be predicted by a pre-trained model; The RNA secondary structure feature is the unpairing probability value of each nucleotide in the RNA sequence to be predicted.

9. An RNA G-quadruplex prediction system based on a pre-trained model and RNA secondary structure, used to implement the prediction method of claim 8, characterized in that: Includes prediction module and feature extraction module; The prediction module is used to output a prediction of whether the RNA sequence to be predicted contains rG4 based on the input RNA sequence data to be predicted; The RNA sequence data includes sequence features and RNA secondary structure features of the RNA sequence to be predicted; The feature extraction module includes a pre-training module and an RNA secondary structure acquisition module. The pre-training module is used to encode the RNA sequence to be predicted to obtain the sequence feature, and the RNA secondary structure acquisition module is used to extract the unpairing probability value of each nucleotide in the RNA sequence to be predicted to obtain the RNA secondary structure feature.

Citation Information

Patent Citations

  • Method and system for predicting target spot of mRNA knock-down by siRNA

    CN113066527A

  • G-quadruplex prediction method based on DNABERT fine tuning

    CN118447929A