A method for predicting lncRNA-miRNA interactions based on sequence complementary site information

By constructing a positive and negative sample set and using Miranda tools to screen binding site information, combined with the use of machine learning models, the problem of difficult to predict lncRNA-miRNA interaction in the prior art is solved, and high accuracy and high efficiency prediction effects are achieved.

CN119207555BActive Publication Date: 2025-06-17YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411092629.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2025-06-17
Estimated Expiration
2044-08-09

AI Technical Summary

Technical Problem

The prior art is difficult to predict potential interactions of isolated lncRNAs or miRNAs that have no interaction with any miRNA or lncRNA, and the interaction spectrum-based approach limits the study of unverified lncRNA-miRNA pairs due to high experimental cost and incompleteness.

Method used

By constructing a positive and negative sample set, the binding site information was extracted using the Miranda tool, the maximum binding score value was screened as representative binding sites, and the k-mer feature, binding score and binding site free energy were extracted as characteristics, and the machine learning model was used to predict lncRNA-miRNA interactions.

Benefits of technology

The accuracy and efficiency of predicting lncRNA-miRNA interactions are improved, the generalization ability and feature expression ability of the model are enhanced, and the prediction success rate and AUC value are significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119207555B_ABST
    Figure CN119207555B_ABST
Patent Text Reader

Abstract

The present invention provides a method for predicting lncRNA-miRNA interactions based on sequence complementary site information, which solves problems such as predicting lncRNA-miRNA interactions. The method includes the following steps: S1: Construction of positive and negative samples; S2: Using Miranda to extract the binding site information of the positive set and the negative set; S3: Selecting the maximum binding score in the positive set and the negative set as the representative binding site representing the interaction between lncRNA and miRNA; S4: Constructing features and feature selection; S5: Using a machine learning model to predict the interaction; S6: Conducting data comparison to evaluate the prediction accuracy of the model. The present invention has the advantages of good prediction effect and high versatility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bioinformatics, and particularly relates to a method for predicting lncRNA-miRNA interactions based on sequence complementary site information. Background Art

[0002] For the method based on graph representation learning, it is difficult for existing methods to predict the potential interactions of isolated lncRNAs or miRNAs that do not interact with any miRNA or lncRNA. Since the sequences of lncRNAs are less conserved among different species, the research conclusions in one species may not be applicable to other species. In the method based on interaction spectra, due to the time-consuming and high-cost of traditional wet experimental methods, which are not suitable for large-scale screening, existing methods rely on data obtained through wet experiments, and these data may be incomplete, limiting the research on unvalidated lncRNA-miRNA pairs. Existing methods based on data mining and deep learning may have limitations in the accuracy and efficiency of predicting lncRNA-miRNA interactions. Existing methods cannot study the relevance of unvalidated lncRNA-miRNA data, and these data may have important research value.

[0003] To solve the deficiencies of the existing technology, people have conducted long-term explorations and proposed various solutions. For example, a Chinese patent document discloses a non-coding RNA interaction prediction model and method based on graph contrast learning [202311482172.2], which first obtains a graph feature extraction module and a prediction module. The graph feature extraction module is used to obtain graph-enhanced data of the lncRNA and miRNA interaction network by means of graph contrast learning; the prediction module predicts the probability of the existence of an interaction between lncRNA and miRNA based on the traversal set of node representations from lncRNA and node representations from miRNA in the graph-enhanced data; both the graph feature extraction module and the prediction module use neural networks.

[0004] The above solution solves the problem of RNA interaction prediction to a certain extent, but there are still many deficiencies in this solution, such as limitations in predicting lncRNA-miRNA interactions. Summary of the Invention

[0005] The object of the present invention is to provide a method for predicting lncRNA-miRNA interactions based on sequence complementary site information, which has good applicability and high accuracy in predicting lncRNA-miRNA interactions, in view of the above problems.

[0006] To achieve the above object, the present invention adopts the following technical solutions: A method for predicting lncRNA-miRNA interactions based on sequence complementary site information, comprising the following steps:

[0007] S1: Construction of positive and negative samples;

[0008] S2: Extract binding site information of the positive set and the negative set using Miranda;

[0009] S3: Select the maximum binding score in the positive set and the negative set as the representative binding site representing the interaction between lncRNA and miRNA;

[0010] S4: Construct features and perform feature selection;

[0011] S5: Use a machine learning model to predict interactions;

[0012] S6: Compare data to evaluate the prediction accuracy of the model.

[0013] In the above method for predicting lncRNA-miRNA interactions based on sequence complementary site information, in step S1, a verified positive sample data set is obtained from the lncRNASNP2 database. The lncRNASNP2 database is selected because it contains experimentally verified lncRNA-miRNA interaction data, ensuring the accuracy and reliability of the samples.

[0014] In the above method for predicting lncRNA-miRNA interactions based on sequence complementary site information, in step S1, human lncRNA sequences are obtained from GENCODE, and human miRNA sequences are extracted from the miRbase database.

[0015] In the above method for predicting lncRNA-miRNA interactions based on sequence complementary site information, the construction of negative samples in step S1 includes the following steps:

[0016] S11: Construct sequence sets with a similarity of greater than or equal to 80% for each lncRNA and miRNA, denoted as and ;

[0017] where Lp represents the sequence set with a similarity of greater than or equal to 80% to any lncRNA sequence p, represents any sequence in the sequence set, Mq represents the sequence set with a similarity of greater than or equal to 80% to any miRNA sequence q, represents any sequence in the sequence set;

[0018] S12: Randomly select one sequence each from lncRNA and miRNA, and label them as Li and Mj respectively. The pair is (L i , M j );

[0019] S13: Extract all miRNAs that interact with L i from the positive samples and denote them as LM i = {m q}. For each m q , extract similar sequences from Mq to form a new similarity matrix . If LM1 i contains M j , then it is considered that there is a potential interaction. Discard the (L i , M j ) interaction pair and repeat step S12. If it does not contain, then retain the (L i , M j ) interaction pair;

[0020] S14: Extract all lncRNAs that interact with M j from the positive samples and denote them as ML j = {l p}. For each l p , extract similar sequences from l p to form a new similarity matrix . If ML1 j contains L i , then it is considered that there is a potential interaction. Discard the (L i , M j ) interaction pair and repeat steps S12 and S13. If it does not contain, then retain the (L i , M j ) interaction pair;

[0021] S15: Repeat steps S12 to S14 until the data volumes of the positive and negative samples reach balance.

[0022] In the above method for predicting lncRNA-miRNA interactions based on sequence complementary site information, step S2 includes the following steps:

[0023] S21: Use the lncRNA and miRNA involved in the interaction as the input of Miranda;

[0024] S22: The Miranda tool contains two key parameters: score and energy. When setting the parameter score to 80 and energy to -1, the generated binding site information can cover all positive samples. We name the set of binding site data obtained under this parameter setting as J;

[0025] S23: Select the binding site information involved in positive and negative samples from J, name the data set of the binding site information extracted from positive samples as Z, and name the data set of the binding site information extracted from negative samples as F;

[0026] S24: In Z and F, each pair of lncRNA-miRNA interactions contains N binding sites, where N is greater than 1. We will screen each site in these interaction pairs, and finally only retain the binding sites that can truly reflect the interaction between lncRNA and miRNA for each pair of interactions;

[0027] S25: Use the set of binding site information obtained after screening the positive set as J p , use the set of binding site information obtained after screening the negative set as J n , and perform the above processing on the binding site information of the negative set as well. After that, extract the binding site sequences of lncRNA and miRNA in J p and J n .

[0028] In the above method for predicting lncRNA-miRNA interactions based on sequence complementary site information, the binding site information set screened by the maximum binding score is obtained in step S3. Name the set of the positive set as J ps , name the set of the negative set as J ns , and after that, extract the binding site sequences of lncRNA and miRNA in J ps and J ns .

[0029] In the above method for predicting lncRNA-miRNA interactions based on sequence complementary site information, the feature selection process in step S4 is in ifeature. Use the sequence in step S3 as the input, and use the k-mer features of the set site sequences and the binding scores and free energies of the binding sites in J ps and J ns as the features of the interaction between lncRNA and miRNA.

[0030] In the above method for predicting lncRNA-miRNA interactions based on sequence complementary site information, step S5 inputs the features extracted in step S4 into a machine learning model, and the machine learning model includes RF, XGBoost, SVM, Decisiontree, LightGBM.

[0031] In the above method for predicting lncRNA-miRNA interactions based on sequence complementary site information, step S5 selects the RF model as the prediction model.

[0032] In the above method for predicting lncRNA-miRNA interactions based on sequence complementary site information, step S6 is compared with the LncMirNet method on the same dataset.

[0033] Compared with the existing technologies, the advantages of the present invention are as follows: taking the binding sites between lncRNA and miRNA as the key factors for predicting their interactions, a balanced positive and negative sample set is constructed through a series of filtering and similarity analyses, which helps to improve the generalization ability and accuracy of the prediction model; after predicting the binding sites using the Miranda tool, through a refined screening process, the most representative binding site information is retained, which helps to improve the prediction accuracy; key features including kmer features, binding scores, and binding site free energies are extracted, and these features are further optimized through a feature selection process, enhancing the feature expression ability of the model; compared with the existing method LncMirNet, there are significant improvements in both the prediction success rate (ACC) and the AUC value, verifying the effectiveness and superiority of the model. Brief Description of the Drawings

[0034] Figure 1 is the flowchart of the method of the present invention;

[0035] Figure 2 is the comparison table of various evaluation indexes between the present invention and the LncMirNet method. Detailed Embodiments

[0036] The present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0037] As Figure 1-2 shown, a method for predicting lncRNA-miRNA interactions based on sequence complementary site information uses the interaction sites between lncRNA and miRNA to predict potential interactions, and it is the first method for predicting lncRNA and miRNA interactions to use binding sites to predict their interactions. Its theoretical basis is that the interaction between lncRNA and miRNA may depend on specific binding sites, and these sites may have specific patterns or motifs in the sequence of lncRNA and interact with the seed region of miRNA (usually the first 2-8 nucleotides of miRNA).

[0038] This method uses Miranda to find the interaction sites between related lncRNAs and miRNAs. There are two parameters in Miranda (score, energy). When the parameter score = 80 and energy = -1 are set, the generated binding sites can cover all positive samples. The data set of binding sites generated under this parameter is named J. So in the subsequent work, the binding site information involved in positive and negative samples is selected from J. Then the data set of binding site information extracted from positive samples is named Z. At this time, each pair of interactions in Z contains N (N>1) binding sites. Then, the N binding pairs contained in each interaction pair are screened. Finally, only the binding sites that can truly reflect the interaction between lncRNA and miRNA are retained for each pair of interactions. Finally, the data set of binding site information obtained after screening the positive set is named J p 。

[0039] The binding site information of the negative set is also processed as above. Next, the binding site sequences of lncRNAs and miRNAs in J p are extracted and feature extraction is performed. The features used in this method are the 2mer features of the sequence, binding score, binding site free energy, and binding site accessibility score. The negative set also performs this step. Finally, these features are used as inputs in the machine learning model to obtain the final prediction results, which specifically include the following steps:

[0040] S1: Construction of positive and negative samples;

[0041] S2: Use Miranda to extract the binding site information of the positive set and the negative set;

[0042] S3: Select the maximum binding score in the positive set and the negative set as the representative binding site representing the interaction between lncRNA and miRNA;

[0043] S4: Perform feature construction and feature selection;

[0044] S5: Use the machine learning model to predict the interaction;

[0045] S6: Perform data comparison to evaluate the prediction accuracy of the model.

[0046] Specifically, in step S1, the positive data set is obtained from the lncRNASNP2 database, and there are real interactions confirmed by laboratory tests and research literature in this database.

[0047] Specifically, in step S1, human lncRNA sequences were obtained from GENCODE, and human miRNA sequences were extracted from the miRbase database. To filter out true interactions, positive lncRNA-miRNA pairs were selected when records in lncRNASNP2 contained both hsa-miR and ENST simultaneously. Finally, 266 miRNAs, 1663 lncRNAs, and 15,640 validated lncRNA-miRNA interactions and their corresponding sequences were obtained.

[0048] Furthermore, the construction of negative samples in step S1 includes the following steps:

[0049] S11: Construct sequence sets with a similarity of 80% or more for each lncRNA and miRNA, denoted as and ;

[0050] where Lp represents the sequence set with a similarity of 80% or more to any lncRNA sequence p, represents any sequence in the sequence set, and m can range from 1 to 1663; Mq represents the sequence set with a similarity of 80% or more to any miRNA sequence q, represents any sequence in the sequence set, and n can range from 1 to 266;

[0051] S12: Randomly select one sequence each from 1663 lncRNAs and 266 miRNAs, marked as Li and Mj respectively, and the combination pair is (L i , M j );

[0052] S13: Extract all miRNAs that interact with L i from the positive samples and denote them as LM i = {m q}, extract similar sequences from Mq for each m q and form a new similarity matrix . If LM1 i contains M j , it is considered that there is a potential interaction, discard the (L i , M j ) interaction pair and repeat step S12; if not, retain the (L i , M j ) interaction pair;

[0053] S14: Extract all lncRNAs that interact with M j from the positive samples and denote them as ML j = {l p}, for each lp Extract similar sequences from l p and construct a new similarity matrix , if ML1 j contains L i then a potential interaction is considered to exist, discard the (L i , M j ) interaction pair and repeat steps S12 and S13; if not, retain the (L i , M j ) interaction pair;

[0054] S15: Repeat steps S12 to S14 until the data volumes of positive and negative samples reach balance.

[0055] In addition, step S2 includes the following steps:

[0056] S21: Use the 1663 lncRNAs and 266 miRNAs involved in the interaction as the input of Miranda;

[0057] S22: The Miranda tool contains two key parameters: score and energy. When setting the parameter score to 80 and energy to -1, the generated binding site information can cover all positive samples. We name the set of binding site data obtained under this parameter setting as J;

[0058] S23: Select the binding site information involved in positive and negative samples from J, take the set of binding site information data extracted from positive samples as Z, and take the set of binding site information data extracted from negative samples as F;

[0059] S24: In Z and F, each pair of lncRNA-miRNA interactions contains N binding sites, where N is greater than 1. We will screen each site in these interaction pairs, and finally only retain the binding sites that can truly reflect the interaction between lncRNA and miRNA for each pair of interactions;

[0060] S25: Take the set of binding site information obtained after screening the positive set as J p , take the set of binding site information obtained after screening the negative set as J n , and perform the above processing on the binding site information of the negative set as well. After that, extract the binding site sequences of lncRNA and miRNA in J p and J n .

[0061] Meanwhile, the set of binding site information selected with the maximum binding score is obtained in step S3. In the interaction between lncRNA and miRNA, a perfect complementary pairing between the seed region of miRNA (usually the first 8 nucleotides) and the target lncRNA will obtain a higher score. When miRNA binds to the target lncRNA, the lower the free energy of the entire double-stranded structure, the more stable the binding, and thus a higher score will be obtained. Some studies have shown that miRNA tends to bind in the 3' untranslated region (3'UTR) of lncRNA, and the binding sites of Miranda in these regions may obtain higher scores. Therefore, in J p and J n selecting the binding site information with the maximum binding score can better represent the interaction characteristics of the entire interaction pair. Name the set of the positive set as j ps , and name the set of the negative set as J ns . After that, extract the binding site sequences of lncRNA and miRNA in J ps and J ns .

[0062] Visibly, in ifeature, the feature selection process of step S4 uses the sequences in step S3 as input, and a total of 24 features regarding rna sequences are tried (2mer, 3mer, ASDC, CKSNAP, DAC, DACC, DCC, DPCP, Geary, mismatch, MMI, Moran, NAC, NMBroto, PCPseDNC, PseDNC(2_0.1), PseKNC(2_0.1_2), PseKNC(2_0.1_3), SCPseDNC, Subsequence, z_curve_9bit, z_curve_12bit, z_curve_36bit, z_curve_48bit). In addition to these features, the binding score and the free energy of the binding site in J ps and J ns are also extracted as the features of the interaction between lncRNA and miRNA. Screen the k-mer features in the sequences and the binding score and the free energy of the binding site in J ps and J ns as the features of the interaction between lncRNA and miRNA.

[0063] Finally, after screening, select the k-mer features of the sequences here and J ps and J nsThe free energy of binding fraction and binding sites, where k-mer is the most basic feature of the RNA sequence, and k-mer is a sequence fragment of length k. For an RNA sequence, a k-mer can be a nucleotide sequence of length k, including four types: adenine (A), cytosine (C), guanine (G), and uracil (U). The k-mer feature is to count the frequency of each k-mer appearing in the RNA sequence and study the distribution pattern of k-mers in the whole sequence. The binding fraction selected in step S3 is the maximum value. The higher this fraction, the more likely the binding between miRNA and its target mRNA will occur and it may be biologically significant. The free energy feature is an important parameter for evaluating the stability of intermolecular interactions, especially in the binding process of biomolecules such as proteins and nucleic acids. In the binding of miRNA to its target lncRNA, the binding free energy is used to quantify the stability of this interaction. The binding free energy refers to the energy released or absorbed when a molecule binds from a free state to another molecule under constant temperature and pressure conditions. For miRNA and its target lncRNA, the binding free energy reflects the stability of their double-stranded structure formation. The binding free energy is usually estimated by calculating the free energy change (ΔG) of the double-stranded RNA (the double-stranded formed by miRNA and its target mRNA). The lower this value (the larger the negative value), the more stable the double-stranded structure and the tighter the binding between miRNA and its target mRNA.

[0064] Obviously, in step S5, the features extracted in step S4 are input into the machine learning model, and the machine learning model includes RF, XGBoost, SVM, Decision tree, and LightGBM. Finally, it is found that the data performs well on RF. Finally, the ACC of the training set on RF reaches 86.334%, and the ACC of the test set reaches 87.02%. Therefore, in step S5, the RF model is selected as the prediction model.

[0065] Obviously, in step S6, a comparison is made with the LncMirNet method on the same dataset. The results show that the prediction success rate has increased by 34.109% in terms of ACC and by 30.96% in terms of AUC value compared with the LncMirNet method. Therefore, these results show that the method for predicting lncRNA-miRNA interactions based on sequence complementary site information has good performance.

[0066] In summary, the principle of this example lies in: innovatively using the interaction sites between long non-coding RNA (lncRNA) and microRNA (miRNA) to predict their potential relationships, and for the first time taking the binding sites as the key factors for predicting their interactions. Experimentally verified positive sample data were collected from the lncRNASNP2 database, and combined with the sequence information in the GENCODE and miRbase databases. A balanced positive and negative sample set was constructed through a series of filtering and similarity analyses. Using the Miranda tool, the binding sites between lncRNA and miRNA were predicted, and the most representative binding site information was screened out from them. Further, key features including kmer features, binding scores, and binding site free energies were extracted, and these features were optimized through a feature selection process. In the comparison of multiple machine learning models, the random forest (RF) model was selected as the final prediction model due to its excellent performance, achieving an accuracy of 86.334% on the training set and 87.02% on the test set.

[0067] The specific examples described in this document are only used to illustrate the principles and applications of the present invention. Those skilled in the art to which the present invention pertains can make various modifications or supplements to the described specific embodiments or use similar ways to substitute, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

[0068] Although terms such as positive and negative samples are used more frequently in this article, the possibility of using other terms is not excluded. The use of these terms is only to more conveniently describe and explain the essence of the present invention; interpreting them as any additional limitation is contrary to the spirit of the present invention.

Claims

1. A method for predicting lncRNA-miRNA interactions based on sequence complementary site information, characterized in that: The steps include: S1: Construction of positive and negative samples; S11: Construct a sequence set with a similarity of 80% or more for each lncRNA and miRNA, denoted as and ; Among them, Lp represents the sequence set with a similarity greater than or equal to 80% with any lncRNA sequence p. represents any sequence in the sequence set, Represents a set of sequences with a similarity greater than or equal to 80% with any miRNA sequence q. Represents any sequence in the sequence set; S12: Randomly select a sequence from lncRNA and miRNA, respectively, and mark them as and , the combination pair is ; S13: Extraction of positive samples and All miRNAs that interact are denoted as , for each from Extract similar sequences and construct a new similarity matrix ,if Include It is considered that there is a potential interaction and discarded Interact with each other and repeat step S12; if it does not contain Interaction pairs; S14: Extract and All lncRNAs that interact are denoted as , for each from Extract similar sequences and construct a new similarity matrix ,if Include It is considered that there is a potential interaction and discarded Interact with each other and repeat steps S12 and S13; if it does not contain Interaction pairs; S15: Repeat steps S12 to S14 until the amount of data of positive and negative samples reaches a balance; S2: Use Miranda to extract the binding site information of positive and negative samples; S21: lncRNA and miRNA involved in the interaction are used as inputs of Miranda; S22: The Miranda tool contains two key parameters: score and energy. When the score is set to 80 and the energy is -1, the generated binding site information can cover all positive samples. The binding site data set obtained under this parameter setting is named J. S23: Select the binding site information involved in the positive and negative samples from J, combine the binding site information data extracted from the positive samples into Z, and combine the binding site information data extracted from the negative samples into F; S24: In Z and F, each pair of lncRNA-miRNA interactions contains N binding sites, where N is greater than 1. Each site in these interaction pairs will be screened, and finally only the binding sites that can truly reflect the interaction between lncRNA and miRNA will be retained for each pair of interactions; S25: The binding site information set obtained after positive sample screening is used as , the binding site information set obtained after negative sample screening is used as The binding site information of negative samples is also processed as above, and then and The binding site sequences of lncRNA and miRNA were extracted; S3: Select the maximum binding score in the positive sample and the negative sample as the representative binding site representing the interaction between lncRNA and miRNA; Step S3 obtains the binding site information set screened by the maximum binding score, and the set of positive samples is named , name the set of negative samples , and then and The binding site sequences of lncRNA and miRNA were extracted; S4: construct features and select features; the feature selection process of step S4 takes the sequence in step S3 as input in ifeature, and selects the k-mer features and and Binding fraction and free energy of binding sites as features of the interaction between lncRNA and miRNA; S5: Predict interactions using machine learning models; S6: Perform data comparison to evaluate the prediction accuracy of the model.

2. A method for predicting lncRNA-miRNA interactions based on sequence complementary site information according to claim 1, characterized in that: The step S1 obtains a verified positive sample data set from the lncRNASNP2 database.

3. The method for predicting lncRNA-miRNA interaction based on sequence complementary site information according to claim 1, characterized in that: The step S1 obtains human lncRNA sequences from GENCODE and extracts human miRNA sequences from the miRbase database.

4. The method for predicting lncRNA-miRNA interaction based on sequence complementary site information according to claim 1, characterized in that: The step S5 inputs the features extracted in step S4 into the machine learning model, and the machine learning model includes RF, XGBoost, SVM, Decision tree, and LightGBM.

5. The method for predicting lncRNA-miRNA interaction based on sequence complementary site information according to claim 4, characterized in that: The step S5 selects the RF model as the prediction model.

6. The method for predicting lncRNA-miRNA interaction based on sequence complementary site information according to claim 4, characterized in that: The step S6 is compared with the LncMirNet method on the same data set.

Citation Information

Patent Citations

  • Non-coding RNA interaction prediction model and method based on graph contrast learning

    CN117497045A

  • Method and system for constructing models for predicting protein-RNA interaction binding sites

    CN111192631A

  • Prediction method for miRNA (micro ribonucleic acid)-lncRNA (long non-coding ribonucleic acid) interaction relationship based on hierarchical deep learning

    CN112270958A

  • Method for predicting interaction relationship between circRNA and miRNA based on ensemble learning

    CN113344076A