A Prediction Method for 2'-O-Methylation Sites in RNA Based on Ensemble Learning

Through the integrated learning method, combined with manual design features and Word2Vec embedding features, the Promoter-BERT model is used for fine-tuning, and the prediction probability is integrated through the soft voting mechanism, the problems of limited prediction accuracy and insufficient generalization ability in RNA were solved, achieving efficient and accurate prediction effects.

CN119207581BActive Publication Date: 2025-07-01ANHUI AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411248190.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-06
Publication Date
2025-07-01
Estimated Expiration
2044-09-06

AI Technical Summary

Technical Problem

In the prior art, the accuracy of 2OM site prediction in RNA is limited, lacks sufficient generalization capabilities, and is complex in operation and expensive.

Method used

Using an integrated learning method, the manual design features such as K-mer, ANF, NCP and Word2Vec embedded features are combined with the Promoter-BERT model for fine-tuning, and the basic model is built with lightweight gradient enhancement institutions, and the predicted probability of each model is integrated through a soft voting mechanism.

Benefits of technology

It improves the accuracy and generalization ability of 2OM site prediction in RNA, reduces operational complexity and cost, and provides an efficient and accurate prediction method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119207581B_ABST
    Figure CN119207581B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for predicting 2'-O-methylation sites in RNA based on ensemble learning, belonging to the technical field of bioinformatics. In view of the characteristics of RNA sequences, the present invention performs task-specific fine-tuning on the Promoter-BERT model, enabling it to more effectively capture the complex patterns of RNA sequences in specific tasks, thereby obtaining high-quality biological feature representations; the ANOVA technique is used to select the extracted features, eliminating redundant features and retaining the most influential features. In addition, traditional sequence features are combined with the embedded features obtained through the Word2Vec model to enhance the expressive power of the model; the prediction results of the lightweight gradient boosting machine and the deep learning model are combined, and a final prediction model is formed through a soft voting mechanism. This ensemble method not only improves the generalization ability of the model but also increases the stability of the prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioinformatics, and particularly relates to a method for predicting 2'-O-methylation (2OM) sites in RNA based on ensemble learning. Background Art

[0002] 2'-O-methylation (2OM) is an important post-transcriptional modification of RNA, which exists in various RNAs such as rRNA, tRNA, mRNA, snoRNA, miRNA, and piRNA. 2OM can stabilize the secondary structure of RNA, especially the helical structure, making it more stable and contributing to maintaining the three-dimensional conformation of RNA molecules. 2OM can also regulate the interaction between RNA and proteins or other RNA molecules, affecting cell signal transduction, gene expression, and regulation processes. In addition, 2OM modification can help the innate immune system distinguish endogenous and exogenous mRNAs, and some viruses evade immune recognition through 2OM modification. Identifying 2OM sites is of great significance for revealing its functions in biology and disease treatment.

[0003] Early experimental techniques such as HClO4 hydrolysis and high-throughput techniques such as Nm-seq have been used for the identification of 2OM sites. However, these methods have their respective limitations. For example, early experimental techniques are complex to operate and may cause irreversible damage to samples, while high-throughput techniques are costly, limiting their application in general laboratories. Computational methods, especially those based on artificial intelligence, can efficiently predict 2OM sites. Some researchers have developed a computational method using support vector machine (SVM), some researchers have constructed the iRNA-2OM model based on SVM, some researchers have developed the iRNA-PseKNC(2methyl) model using convolutional neural network (CNN), and some researchers have developed an ensemble model by combining seven optimal classifiers. In addition, there are also prediction models such as NmSEER and DeepOMe developed based on other datasets.

[0004] Traditional 2OM site prediction models rely on manually designed features such as Nucleotide Composition and secondary structure prediction. These features may not be able to comprehensively capture the complex biological properties of RNA sequences, resulting in limited prediction accuracy. Existing computational models, such as support vector machine (SVM) and basic neural networks, are often trained on small-scale or single datasets. These models may perform poorly on new or diverse datasets and lack sufficient generalization ability. Many high-throughput experimental methods, such as Nm-seq, although able to provide accurate modification site information, are complex to operate and costly, and are not suitable for large-scale or rapid biological research needs.

[0005] The above problems need to be solved urgently. For this purpose, the present invention proposes a method for predicting 2OM sites in RNA based on ensemble learning. Summary of the Invention

[0006] The technical problem to be solved by the present invention is: how to solve the problems of limited prediction accuracy, lack of sufficient generalization ability, complex operation and high cost existing in the prior art, and provides a method for predicting 2OM sites in RNA based on ensemble learning.

[0007] The present invention solves the above technical problems through the following technical solutions. The present invention includes the following steps:

[0008] S1: Data Preparation

[0009] Collect 2OM site data from public databases and randomly divide it into a training set and a test set according to a set ratio.

[0010] S2: Feature Learning

[0011] Extract three manually designed features from the RNA sequence using K-mer, ANF, and NCP, and perform vectorization processing on the RNA sequence using the Word2Vec model to obtain Word2Vec embedding features; use ANOVA technology to select a feature subset highly relevant to 2OM site prediction from the manually designed features and combine it with the embedding features generated by the Word2Vec model.

[0012] S3: Model Establishment

[0013] Based on 1mer, 3mer, and 5mer word segmentation to obtain texts of sequence units of different lengths, that is, segmented sequences, and input the segmented sequences into a pre-trained Promoter-BERT model for fine-tuning to obtain three fine-tuned models; based on the obtained feature subset, use a lightweight gradient boosting mechanism to build a base model, and integrate it with some models obtained by fine-tuning the pre-trained Promoter-BERT model. Integrate the prediction probabilities of each base model through a soft voting mechanism to form the final integrated model 2OMPro, where some models obtained by fine-tuning the pre-trained Promoter-BERT model are also base models.

[0014] S4: Optimization and Evaluation

[0015] Use cross-validation and an independent test set to evaluate the performance of the integrated model 2OMPro.

[0016] S5: Deployment and Prediction

[0017] Deploy the integrated model 2OMPro in actual biomedical research work to predict unknown 2OM sites.

[0018] Furthermore, in the step S1, the specific processing procedure is as follows:

[0019] S11: Collect the original 2OM site data from the RMBase v2.0 and the experimental dataset generated based on the Nm-seq technology;

[0020] S12: For each 2OM site data, cut out RNA fragments with a length of a set number of nucleotides from the upstream and downstream of the 2OM site as positive samples; at the same time, randomly select sequences with the same length from the upstream and downstream of each 2OM site as negative samples;

[0021] S13: Set the sequence similarity threshold, and use the CD-HIT software to perform sequence clustering to remove redundant sequences, obtaining a dataset with equal positive and negative samples, that is, obtaining an RNA sequence dataset;

[0022] S14: Based on the modified base types of the 2OM sites, divide the dataset into four subsets, and randomly divide each subset into a training set and a test set according to a set ratio.

[0023] Furthermore, in the step S2, in order to cover sequence information at different levels, the RNA sequences are segmented into 1mer, 3mer, and 5mer as vocabulary, and then the RNA sequences are input into the pre-trained Promoter-BERT model as sentences for fine-tuning, obtaining a fine-tuned Promoter-BERT model based on 1mer, 3mer, and 5mer.

[0024] Furthermore, in the step S2, the expression for the fine-tuned Promoter-BERT model to encode the RNA sequences is as follows:

[0025] X BERT = BERT(S)

[0026] where S represents the input RNA sequence, BERT represents the fine-tuned Promoter-BERT model, and X BERT represents the encoding of the input sequence S by the fine-tuned Promoter-BERT model.

[0027] Furthermore, in the step S2, the three manually designed features are nucleotide composition K-mer, nucleotide compound property NCP, and auto-correlation property ANF; among them, the nucleotide composition K-mer is used to calculate the frequency of consecutive k nucleotides to extract the short-range information of the RNA sequence, the nucleotide compound property NCP is used to represent A, U, C, and G by vector encoding (1, 1, 1), (0, 0, 1), (0, 1, 0), and (1, 0, 0) respectively, and the auto-correlation property ANF is used to describe the distribution of nucleotides in the RNA sequence and their positional correlation in the sequence by calculating the cumulative occurrence frequency d of each nucleotide from the start of the sequence to position i i is achieved by

[0028] Furthermore, in the nucleotide composition K-mer, using the k-mer frequency, each primary RNA sequence R can be transformed into a vector with 4k elements as follows:

[0029]

[0030] where is the normalized occurrence frequency of the i-th k-tuple:

[0031]

[0032] where, n i is the number of the i-th k-tuple nucleotide components in the RNA sequence, and L represents the length of the RNA sequence

[0033] Furthermore, in the auto-correlation property ANF, the cumulative occurrence frequency d of each nucleotide from the start of the sequence to position i in the RNA sequence i is defined as:

[0034]

[0035] where, |N i | represents the length of the prefix subsequence from the start of the sequence to position i, that is, the number of nucleotides before the i-th position, and f(n j ) is an indicator function, indicating that when the nucleotide n j at position j is the same as the target nucleotide ni, f(n j ) is equal to 1, otherwise 0

[0036] Furthermore, in the step S2, the specific processing procedure for selecting a feature subset highly correlated with the 2OM site prediction from the manually designed features using the ANOVA technique is as follows:

[0037] Suppose there are m manually designed features, then the feature subset F after screening by the ANOVA techniqueselected Denoted as:

[0038] F selected ={f i ∣p(f i )≤α}

[0039] where f i represents the i-th manually designed feature, p(f i ) represents the p-value of the ANOVA test for the i-th feature, and α represents the set significance level for feature screening.

[0040] Furthermore, in the step S3, the partial models obtained by fine-tuning the pre-trained Promoter-BERT model as the base model are respectively the fine-tuned Promoter-BERT models based on 3mer and 5mer.

[0041] Furthermore, in the step S3, if there are M base models and N classes, the prediction probability of the m-th base model for the n-th class is denoted as Pmn, then the prediction probability P soft (n) of the n-th class under soft voting is calculated as:

[0042]

[0043] where P mn represents the prediction probability of the m-th base model for the n-th class;

[0044] The soft voting mechanism averages the prediction probabilities of all base models for each class, and the final result C is the final prediction result of the integrated model 2OMPro by the class with the highest average prediction probability:

[0045] C = argmax n P soft (n)

[0046] where argmax n means taking the n value that makes P soft (n) the largest, and n is the class number.

[0047] The present invention has the following advantages compared with the prior art:

[0048] 1. For the characteristics of RNA sequences, the Promoter-BERT model is fine-tuned for specific tasks, enabling it to more effectively capture the complex patterns of RNA sequences in specific tasks, thereby obtaining high-quality biological feature representations;

[0049] 2. Select the extracted features using ANOVA technology, eliminate redundant features, and retain the most influential features. In addition, combine traditional sequence features with the embedded features obtained through the Word2Vec model to enhance the model's expressive ability;

[0050] 3. Combine the prediction results of the lightweight gradient boosting machine and the deep learning model, and form the final prediction model through a soft voting mechanism. This integration method not only improves the generalization ability of the model but also increases the stability of the prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a schematic flow chart of the present invention;

[0052] Figure 2 is a five-fold cross-validation result diagram of the Promoter-BERT model after fine-tuning on each subset in the embodiment of the present invention, where a is the Am subset, b is the Cm subset, c is the Gm subset, and d is the Um subset;

[0053] Figure 3 is the ROC curve and AUC value of each model under different feature combinations on each subset in the embodiment of the present invention, where a is the Am subset, b is the Cm subset, c is the Gm subset, and d is the Um subset. DETAILED DESCRIPTION OF THE INVENTION

[0054] The following is a detailed description of the embodiments of the present invention. The embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments.

[0055] As Figure 1 shown, this embodiment provides a technical solution: a method for predicting 2OM sites in RNA based on ensemble learning, mainly including the following steps:

[0056] 1. Data Preparation

[0057] Collect 2OM site data from public databases and randomly divide it into a training set and a test set in a ratio of 7:3.

[0058] 2. Feature Learning

[0059] Extract three manually designed features from the RNA sequence using K-mer, ANF, and NCP, and perform vectorization processing on the RNA sequence using the Word2Vec model to obtain Word2Vec embedded features; use ANOVA technology to select a feature subset highly relevant to 2OM site prediction from the manually designed features and combine it with the embedded features generated by the Word2Vec model.

[0060] 3. Model Establishment

[0061] Based on the text of sequence units of different lengths obtained by segmenting with 1mer, 3mer, and 5mer, i.e., the segmented sequences, the segmented sequences are input into the pre-trained Promoter-BERT model for fine-tuning, and three fine-tuned models are obtained; based on the obtained feature subsets, a basic model is constructed using a lightweight gradient boosting mechanism, and it is integrated with the model obtained by fine-tuning the pre-trained Promoter-BERT model. The prediction probabilities of each basic model are integrated through a soft voting mechanism to form the final integrated model 2OMPro.

[0062] S4: Optimization and evaluation

[0063] The performance of the integrated model 2OMPro is evaluated using cross-validation and an independent test set.

[0064] S5: Deployment and prediction

[0065] The integrated model 2OMPro is deployed in actual biomedical research work to predict unknown 2OM sites.

[0066] The above steps are further described as follows:

[0067] Step 1: Data preparation

[0068] To ensure a fair comparison with other advanced models, the present invention selects 7597 original 2OM site data collected from the RMBase v2.0 and experimental datasets generated based on the Nm-seq technology. These data have been stored in the GEO database (GSE90164). These 2OM sites are widely distributed in various regions of RNA, including coding sequences (CDS), 3'-UTR, 5'-UTR, introns, exons, and intergenic regions, covering almost all types of RNA such as tRNA, rRNA, scRNA, scaRNA, and snRNA. To maintain the consistency of data processing, for each 2OM site, RNA fragments of 20 nucleotides in length are cut from the upstream and downstream of the site as positive samples. At the same time, sequences of the same length (41 nucleotides) are randomly selected from the upstream and downstream of each 2OM site (the central position nucleotides of these sequences are not methylated) as negative samples, and the ratio of positive samples to negative samples is approximately 1:3.

[0069] In addition, to reduce data redundancy, the CD-HIT software is used for sequence clustering in the present invention, and the cut-off rate (sequence similarity threshold) is set to 80%. After reprocessing, a dataset with an equal number of positive and negative samples is finally obtained, totaling 6091 samples. Based on the modified base types at the 2OM sites, the dataset is further divided into four subsets: Am, Um, Cm, Gm, (as shown in Table 1) and each subset is randomly divided into a training set and a test set at a ratio of 7:3. To pre-train the BERT model, the UCSC dataset from the UCSC Genome Browser is used in the present invention, which is one of the databases widely used in the biological field.

[0070] Table 1 Details of the dataset

[0071]

[0072] Step 2: Feature learning

[0073] 2.1 Fine-tuning of the Promoter-BERT model:

[0074] In the present invention, first, according to the specific characteristics of RNA sequences, the pre-trained Promoter-BERT model for understanding the language characteristics of DNA sequences is fine-tuned. The Promoter-BERT model is originally based on the Transformer architecture and is pre-trained to understand the language characteristics of DNA sequences. In the present invention, the model is re-trained to adapt to the feature recognition of 2OM sites. To cover different levels of sequence information, three Promoter-BERT models based on 1mer, 3mer, and 5mer vocabularies are fine-tuned respectively, and the weights of the models are adjusted to better adapt to specific biological tasks. Complex relationships from single nucleotides to longer sequence fragments can be captured.

[0075] X BERT = BERT(S)

[0076] where S represents the input RNA sequence, BERT represents the fine-tuned Promoter-BERT model, and X BERT represents the encoding of the input sequence S by the fine-tuned Promoter-BERT model.

[0077] 2.2 Feature extraction and fusion:

[0078] The present invention extracts three manually designed features from RNA sequences: nucleotide composition (K-mer), nucleotide chemical property (NCP), and auto-correlation feature (ANF). These features depict the basic physical and chemical properties of RNA sequences, providing preliminary information about the sequence structure and function for the model. To enhance the model's understanding of the context information of RNA sequences, the present invention also uses the Word2Vec model to vectorize the sequence data. This method can convert sequence information into a dense vector form, thereby helping the model capture and utilize local and global context dependencies in the sequence. Then, through the analysis of variance (ANOVA) technique, the above three manually designed features (K-mer, NCP, and ANF) are screened to select the most informative feature subset. The screening and fusion of these manually designed features enhance the diversity and richness of feature expression and optimize the feature usage efficiency in the model training process. Finally, these selected and fused feature sets, together with the embedding features generated by Word2Vec, constitute the basis for model training, providing strong data support for the highly accurate prediction of 2OM sites.

[0079] Suppose there are m manually designed features (K-mer, NCP, and ANF), then the feature subset F after ANOVA screening selected can be expressed as:

[0080] F selected ={f i ∣p(f i )≤α}

[0081] where f i represents the i-th manually designed feature, p(f i ) represents the p-value of the ANOVA test for the i-th feature, and α represents the set significance level for feature screening.

[0082] 2.2.1 NCP

[0083] According to their properties, A, U, C, and G are divided into multiple categories, as shown in Table 2. This descriptor has been widely used in the prediction of various nucleotide modification sites. Specifically, A and G are bicyclic structures, C and U are monocyclic structures; A and C belong to amino groups, G and U belong to keto groups; A and U form weak hydrogen bonds, and C and G can form strong hydrogen bonds. Therefore, based on these chemical properties, A, U, C, and G can be represented by the coordinates (vector encoding) (1, 1, 1), (0, 0, 1), (0, 1, 0), and (1, 0, 0) respectively.

[0084] Table 2 Classification information and corresponding coordinate representations

[0085]

[0086] 2.2.2ANF

[0087] In the analysis of RNA sequences, the autocorrelation feature (ANF) is used to describe the distribution of nucleotides in a sequence and their position correlation in the sequence by calculating the cumulative frequency d of each nucleotide from the beginning of the sequence to position i. i Specifically, the density d of the nucleotide at position i in the RNA sequence is i is defined as:

[0088]

[0089] Among them, |N i | represents the length of the prefix subsequence from the beginning of the sequence to position i, that is, the number of nucleotides before the i-th position. f(n j ) is an indicator function that indicates when nucleotide n at position j j With target nucleotide n i When the same, f(n j ) is equal to 1, otherwise it is 0.

[0090] 2.2.3 K-mer

[0091] k-mer calculates the frequency of consecutive k nucleotides to extract short-range information of RNA sequences. Using k-mer frequencies, each primary RNA sequence R can be converted into a vector with 4k elements as follows:

[0092]

[0093] in, is the normalized frequency of occurrence of the i-th k-tuple:

[0094]

[0095] Among them, n i is the number of nucleotide components in the i-th k-tuple in the RNA sequence, and L represents the length of the RNA sequence.

[0096] Step 3: Model building

[0097] During the feature encoding process, the Promoter-BERT model has been fine-tuned according to the characteristics of RNA sequences. Figure 2 The five-fold cross validation results of the fine-tuned model are shown. Figure 2 It can be seen that no matter what the data set is, the prediction results of 3mer and 5mer are very similar, and all indicators are much better than 1mer. Therefore, we choose the models of 3mer and 5mer as the basic models.

[0098] Based on the above features, a basic model is constructed using LightGBM (LGBM). To improve the prediction accuracy and stability, the present invention integrates two fine-tuned models based on Promoter-BERT with the feature model based on LGBM, and integrates the prediction probabilities of each model through a soft voting mechanism to improve the prediction stability and accuracy. In soft voting, the prediction probabilities of each basic model are used to calculate the final integrated output. Specifically, the prediction probability of each category is the weighted average of the prediction probabilities of all basic models for that category.

[0099] If there are M basic models and N categories, and the prediction probability of the m-th basic model for the n-th category is denoted as Pmn, then the prediction probability P soft (n) under soft voting is calculated by the formula:

[0100]

[0101] where P mn represents the prediction probability of the m-th basic model for the n-th category;

[0102] The soft voting mechanism averages the prediction probabilities of all basic models for each category, and the final result C is the category with the highest average prediction probability as the final prediction result of the integrated model 2OMPro:

[0103] C = argmax n P soft (n).

[0104] where argmax n means taking the n value that makes P soft (n) the largest, and n is the category number.

[0105] Step 4: Model evaluation and result analysis

[0106] In the experiment, we used metrics such as accuracy (ACC), Matthews correlation coefficient (MCC), precision, and recall to evaluate the performance of the model (2OMPro). Their definitions are as follows:

[0107]

[0108] Among them, TP, TN, FP, and FN represent the numbers of true positives, true negatives, false positives, and false negatives, respectively. The area under the receiver operating characteristic curve (AUROC) is also used, which comprehensively evaluates the performance of the model by balancing the true positive rate (also known as sensitivity) and the false positive rate. Its value ranges from 0 to 1, and the closer it is to 1, the better the model performance. AUROC is not affected by class imbalance and data distribution, so it has strong robustness. AUROC is a valuable metric for comparing the performance of different models and helps to determine the optimal classification threshold that balances sensitivity and specificity. The results are as Figure 3 shown.

[0109] For the Am dataset, the AUC value of the 3mer model is 0.9300, the AUC value of the 5mer model is 0.9325, the AUC value of the 3-5mer combination is 0.9365, the AUC value of the 3mer and LGBM model is 0.9418, the AUC value of the 5mer and LGBM model is 0.9437, and the AUC value of the 3-5mer and LGBM model is 0.9451. This trend is consistent across different datasets (including Cm, Gm, and Um), where the combination of the 3mer and 5mer models with the LGBM model shows the best performance. This improvement is attributed to the combination of the popular large language model BERT and traditional feature extraction methods, enabling the model to effectively capture RNA sequence information and thus obtain more efficient prediction results.

[0110] In the present invention, the integrated model 2OMPro model consists of three key submodels: two base models are fine-tuned from the internal DNA language model Promoter-BERT; the third base model is constructed using the Light Gradient Boosting Machine (LGBM) based on features extracted from RNA sequences and fused with Word2Vec embeddings. These submodels are integrated into an integrated model that combines their prediction probabilities.

[0111] The present invention marks the first application of the large language model (Promoter-BERT) to 2OM site prediction, significantly improving the performance by leveraging its ability to extract subtle information from RNA sequences. The integrated algorithm design integrates the predictions of the three base models, reducing the prediction errors of each model and thus producing more stable and reliable predictions. The results on four different datasets show that the performance of this embodiment is consistently better than that of existing state-of-the-art models, demonstrating its robustness and effectiveness.

[0112] The model of the present invention is not only applicable to human RNA, but can also be extended to the RNA analysis of other species and is suitable for various types of RNA modification research, such as methylation, acetylation, etc. Compared with traditional experimental methods, the present invention provides a cost-effective, fast and accurate prediction method, which is particularly suitable for processing large-scale data and greatly saves time and economic costs.

[0113] In summary, the technical solution of the present invention is innovative in theory and shows remarkable effects and advantages in practical applications. Through precise feature processing and efficient model design, the prediction performance of RNA 2'-O-methylation sites is effectively improved, which helps to promote the research and application development in related biomedical fields.

[0114] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for predicting 2OM sites in RNA based on ensemble learning, characterized in that: The following steps are involved: S1: Data Preparation 2OM site data were collected from public databases and randomly divided into training and test sets according to a set ratio; S2: Feature Learning Three manually designed features, K-mer, ANF and NCP, were extracted from RNA sequences, and the RNA sequences were vectorized using the Word2Vec model to obtain Word2Vec embedded features. ANOVA technology was used to select a subset of features highly correlated with 2OM site prediction from the manually designed features, and combined with the embedded features generated by the Word2Vec model. In step S2, the three manually designed features are nucleotide composition K-mer, nucleotide compound property NCP, and auto-correlation property ANF; wherein the nucleotide composition K-mer is used to calculate the frequency of k consecutive nucleotides to extract the short-range information of the RNA sequence, the nucleotide compound property NCP is used to use vector encoding (1,1,1), (0,0,1), (0,1,0) and (1,0,0) to represent A, U, C and G respectively, and the auto-correlation property ANF is used to describe the distribution of nucleotides in the RNA sequence and their position correlation in the sequence, by calculating the cumulative occurrence frequency d of each nucleotide from the beginning of the sequence to position i i to achieve; S3: Model building Based on 3mer and 5mer word segmentation, texts of sequence units of different lengths are obtained, i.e., segmented sequences. The segmented sequences are input into the pre-trained Promoter-BERT model for fine-tuning, and two fine-tuned models are obtained. Based on the obtained feature subset, a basic model is constructed using a lightweight gradient boosting machine, and it is integrated with the model obtained by fine-tuning the pre-trained Promoter-BERT model. The prediction probabilities of each basic model are integrated through a soft voting mechanism to form the final integrated model 2OMPro, in which the model obtained by fine-tuning the pre-trained Promoter-BERT model is also the basic model. S4: Optimization and Evaluation The performance of the integrated model 2OMPro was evaluated using cross-validation and independent test sets; S5: Deployment and prediction The integrated model 2OMPro is deployed in actual biomedical research to predict unknown 2OM sites.

2. A method for predicting 2OM sites in RNA based on ensemble learning according to claim 1, characterized in that: In step S1, the specific processing process is as follows: S11: Collect raw 2OM loci data from RMBase v2.0 and experimental datasets generated based on Nm-seq technology; S12: For each 2OM site data, RNA fragments with a set length of nucleotides were cut from the upstream and downstream of the 2OM site as positive samples; at the same time, sequences of the same length were randomly selected from the upstream and downstream of each 2OM site as negative samples; S13: Set the sequence similarity threshold, use CD-HIT software to perform sequence clustering to remove redundant sequences, and obtain a data set with equal positive samples and negative samples, that is, obtain an RNA sequence data set; S14: Based on the modified base types of the 2OM site, the dataset was divided into four subsets, and each subset was randomly divided into a training set and a test set according to a set ratio.

3. A method for predicting 2OM sites in RNA based on ensemble learning according to claim 1, characterized in that: In step S3, in order to cover different levels of sequence information, the RNA sequence is segmented into 3mers and 5mers as vocabulary, and then the RNA sequence is input as a sentence into a pre-trained Promoter-BERT model for fine-tuning to obtain a fine-tuned Promoter-BERT model based on 3mers and 5mers.

4. A method for predicting 2OM sites in RNA based on ensemble learning according to claim 1, characterized in that: In step S3, the expression for encoding the RNA sequence by the fine-tuned Promoter-BERT model is as follows: X BERT =BERT(S) Where S represents the input RNA sequence, BERT represents the fine-tuned Promoter-BERT model, and X BERT Represents the encoding of the input sequence S by the fine-tuned Promoter-BERT model.

5. A method for predicting 2OM sites in RNA based on ensemble learning according to claim 1, characterized in that: In the nucleotide composition K-mer, using the k-mer frequency, each primary RNA sequence R can be converted into a vector with 4k elements as follows: in, is the normalized frequency of occurrence of the i-th k-tuple: Among them, n i is the number of nucleotide components in the i-th k-tuple in the RNA sequence, and L represents the length of the RNA sequence.

6. A method for predicting 2OM sites in RNA based on ensemble learning according to claim 1, characterized in that: In the autocorrelation feature ANF, the cumulative occurrence frequency d of each nucleotide in the RNA sequence from the beginning of the sequence to position i i Defined as: Among them, |N i | represents the length of the prefix subsequence from the beginning of the sequence to position i, that is, the number of nucleotides before the i-th position, f(n j ) is an indicator function that indicates when nucleotide n at position j j With target nucleotide n i When the same, f(n j ) is equal to 1, otherwise it is 0.

7. A method for predicting 2OM sites in RNA based on ensemble learning according to claim 1, characterized in that: In step S2, the specific process of selecting a feature subset highly correlated with 2OM site prediction from the manually designed features using the ANOVA technique is as follows: Assuming there are m manually designed features, the feature subset F selected by ANOVA technology selected It is expressed as: F selected ={f i ∣p(f i )≤α} Among them, f i represents the i-th manually designed feature, p(f i ) represents the p-value of the ANOVA test of the ith feature, and α represents the set significance level for feature screening.

8. A method for predicting 2OM sites in RNA based on ensemble learning according to claim 1, characterized in that: In step S3, if there are M basic models and N categories, the predicted probability of the mth basic model for the nth category is denoted as Pmn, then the predicted probability of the nth category under soft voting is P soft (n) The calculation formula is: Among them, P mn Represents the predicted probability of the mth basic model for the nth class; The soft voting mechanism averages the predicted probabilities of all base models for each category, and the final result C consists of the category with the highest average predicted probability as the final prediction result of the integrated model 2OMPro: C=argmax n P soft (n) Among them, argmax n Denotes that P soft (n) The maximum value of n, where n is the category number.

Citation Information

Patent Citations

  • Multi-type RNA methylation modification site prediction method

    CN115273965A

  • Method for predicting acetylation sites of biological lysine

    CN116798514A