MiRNA-drug-side effect triple association prediction method
By integrating multi-source data and using a variational hypergraph encoding and contrastive learning method, and generating enhanced views using a hypergraph convolutional network, the problems of data sparsity and heterogeneity in the prediction of miRNA-drug-side effect triple associations are solved, achieving high-precision and high-generalization prediction.
Patent Information
- Application Number
- CN202511720132.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-02-27
AI Technical Summary
Existing technologies struggle to effectively identify and predict the triplet relationships between miRNAs, drugs, and side effects, especially under conditions of data sparsity and heterogeneity, where traditional methods are unable to capture complex biological processes and high-order topological semantics.
We employ variational hypergraph encoding and contrastive learning methods to construct a high-quality benchmark dataset by integrating multi-source data. We then use a hypergraph convolutional network to learn triple representations and perform end-to-end training using a joint objective function of supervised binary classification loss, contrastive consistency loss, and generator reconstruction loss to generate enhanced views and improve prediction accuracy.
Under conditions of limited annotation and data heterogeneity, the ability to identify miRNA-drug-side effect triple associations was significantly improved, prediction accuracy and generalization performance were enhanced, and a high-quality benchmark dataset was constructed, which can effectively identify complex biological processes and high-order topological dependencies.
Smart Images

Figure CN121583318A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of bioinformatics, and particularly relates to a miRNA-drug-side effect ternary group correlation prediction method. BACKGROUND
[0002] More and more studies show that the recognition of the interaction among miRNA, drug and side effect has important scientific and clinical significance. In drug mechanism and safety evaluation, the off-target effect of miRNA is an important source of drug side effects. Chemical small molecule drugs can directly or indirectly target miRNA, which is an important strategy for treating complex diseases such as cancer. However, a single miRNA can regulate hundreds of target genes, and drugs targeting specific miRNAs may disturb other pathways and cause side effects.
[0003] Miravirsen, an antisense oligonucleotide inhibitor targeting miR-122, is commonly used in the treatment of hepatitis C virus (HCV), but may inhibit the physiological function of miR-122 and cause related adverse reactions. Researchers Bhatt et al. found in BUMPT-306 cell lines and wild-type and p53-deficient mice that cisplatin, commonly used as an anti-tumor drug, can induce up-regulation of miR-34a and is associated with side effects such as nephrotoxicity. In summary, the research on the correlation of the drug-miRNA-side effect (MDS) ternary group is still ongoing, and related research will help to clarify the molecular mechanism of drugs, drug safety evaluation and precise medication.
[0004] At present, the biological experimental techniques for identifying the correlation of MDS ternary group include stable cell line construction, luciferase biosensor assay, Western blot protein detection, RNA immunoprecipitation (RIP), etc. Although these techniques have high recognition accuracy, they generally have the problems of long time-consuming, high labor intensity, high cost and limited throughput. With the advancement of drug research and the continuous accumulation of related data, scientists have carried out correlation prediction research based on calculation. This kind of research is relatively efficient and can carry out large-scale screening and sorting of drugs at a lower cost. The research provides clues for subsequent experimental verification, thereby improving the efficiency of biological network correlation research.
[0005] Existing computational methods in the field of drug discovery focus on binary relationship prediction. Research on the binary relationship between miRNA-drug and drug-side effects has made some achievements, but there is still no computational research tool developed for the prediction of MDS triplets. Among them, the identification of miRNA-drug binary association researches mostly focus on small molecule drugs. Early studies used traditional machine learning or network-based methods, which had limitations such as ignoring topological information and being unable to predict new entity associations. In recent years, scientists have used deep learning models to conduct related binary association research, using graph convolutional neural networks (GNN), deep autoencoders, and other methods. These methods have achieved good results, but they still focus on binary relationships.
[0006] Drug side effect research aims to improve drug safety and is the core problem of drug safety evaluation. Related work based on machine learning is usually modeled as three types of tasks: binary classification, multi-classification, and multi-label classification. Binary classification is used to determine whether a drug has a specific side effect, and representative methods include the GGSC method and the TCSD method. Multi-classification and multi-label models mostly focus on side effects caused by drug-drug interactions (DDI), where multi-classification models predict the most important or most severe side effects of a specific DDI sample, such as MCFF-MTDDI, ACDGNN, and DSN-DDI. Multi-label models address multiple possible side effects. These models generally follow four technical routes: information network analysis based on known associations, traditional machine learning based on drug and side effect features, matrix factorization and embedding learning-based methods, and deep learning. In recent years, there have also been approaches that focus on knowledge graphs and literature and electronic medical record text mining. However, existing methods have not generally modeled and predicted the ternary relationship between miRNA, drugs, and side effects.
[0007] In the computational methods for MDS triplet identification, there are the following challenges. First, the sparsity and heterogeneity of the data are the main challenges. The known association data in MDS data is relatively limited, and the data comes from various sources. How to effectively integrate these heterogeneous data and extract meaningful features from them is the focus of the research. Second, the interaction mechanism between miRNA, drugs, and side effects is complex and involves multiple levels of biological processes. Traditional methods are difficult to capture these deep nonlinear relationships.
[0008] In recent years, the rapid development of deep learning has opened up new paths for complex relationship prediction in bioinformatics, especially the contrastive learning and hypergraph modeling techniques, which have shown outstanding advantages in alleviating data sparsity and robust representation. In this context, this application uses variational hypergraph encoding to automatically generate enhanced views from the original views while preserving high-order topological semantics. By comparing the consistency between the original view and the enhanced view, effective structural signals are mined from unlabeled data. SUMMARY
[0009] To overcome the above deficiencies of the prior art, the present application provides a miRNA-drug-side effect triplet prediction method. To solve the technical problems that it is difficult to unify modeling and mine high-order correlations of the three under the conditions of limited annotation and sparse and heterogeneous data, resulting in insufficient prediction accuracy and generalization ability of MDS triplets. The three joint objective functions composed of supervised binary classification loss, contrast consistency loss and generator variational reconstruction loss are used, which simultaneously considers the discrimination performance and generalization stability. Under the conditions of limited annotation and heterogeneous data, the recognition ability of MDS triplet correlation is significantly improved.
[0010] To solve the above technical problems, the present application adopts the following technical solutions:
[0011] A miRNA-drug-side effect triplet correlation prediction method, comprising the following steps:
[0012] S1: Integrating multiple source data to construct a benchmark data set for miRNA-drug-side effect triplet recognition;
[0013] S2: Based on the benchmark data set in step S1, model the miRNA-drug-side effect correlation as a hypergraph structure;
[0014] S3: Based on the hypergraph structure constructed in step S2, learn the miRNA-drug-side effect triplet representation using a hypergraph convolution network;
[0015] S4: Construct an enhanced view based on a variational hypergraph autoencoder;
[0016] S5: Based on the contrast learning framework, the hypergraph structure and the enhanced view are used as a whole, and the joint loss function composed of supervised binary cross entropy loss, contrast consistency loss and generation reconstruction loss is used for end-to-end training to improve the discrimination ability and generalization performance of the model;
[0017] S6: Use the trained model to predict the correlation of the candidate miRNA-drug-side effect triplet, and output the prediction score.
[0018] Further, in step S1, the integration of multiple source data includes: extracting experimentally verified miRNA-drug correlation data from ncRNA Drug database, collecting experimentally verified drug-side effect correlation data from DrugBank and SIDER database, downloading miRNA sequence data from miRBase, and obtaining drug SMILES string data from DrugBank.
[0019] Further, the construction of the benchmark dataset further comprises: aligning and integrating the miRNA-drug association data and the drug-side effect association data through the drug unique identifier, constructing an initial MDS triad dataset; and further screening and optimizing the initial MDS triad dataset to finally obtain the benchmark dataset of the MDS triad.
[0020] Further, in step S2, the hypergraph structure is defined as G=(V,E), wherein: the node set V is composed of the miRNA set M, the drug set D, and the side effect set S, i.e., V=M∪D∪S; each hyperedge e in the hyperedge set E corresponds to a miRNA-drug-side effect triad (mi,dj,sk) in the benchmark dataset in step S1, and is used to represent the collaborative association of the three.
[0021] Further, in step S2, the hypergraph structure further comprises an association matrix wherein the matrix element represents that the node belongs to the hyperedge , represents that the node does not belong to the hyperedge , and is used to quantify the membership relationship between the node and the hyperedge.
[0022] Further, in step S3, the miRNA-drug-side effect triad representation is learned by using the hypergraph convolution network, which specifically comprises:
[0023] (1) Extracting the feature representations of miRNA, drug, and side effect respectively: calculating the similarity between miRNAs using cosine similarity and mapping the miRNA features through a fully connected network, constructing a molecular graph based on the drug SMILES and encoding the drug features through an isomorphic graph network, and calculating the similarity between side effects using Jaccard coefficient and mapping the side effect features through a fully connected network;
[0024] (2) Concatenate the above features into a hypergraph node attribute matrix, input the hypergraph convolution network, and learn the high-order association features of the nodes through multiple layers of hypergraph convolution operations;
[0025] (3) Concatenate the node features of miRNA, drug, and side effect to obtain the joint representation of the miRNA-drug-side effect triad.
[0026] Further, in step S4, the enhanced view is constructed based on the variational hypergraph autoencoder, which specifically comprises:
[0027] (1) Convert the hypergraph constructed in step S2 into a node-hyperedge equivalent bipartite graph, wherein the vertex set contains entity nodes and hyperedges, and the edge set represents the membership relationship between entity nodes and hyperedges;
[0028] (2) A two-path hypergraph neural network is used as an encoder to learn the latent variable distribution of entity nodes and hyperedges respectively;
[0029] (3) The association probability of entity nodes and hyperedges is reconstructed through a decoder, and a new bipartite graph edge set is generated by combining Gumbel-Softmax differentiable sampling;
[0030] (4) The new bipartite graph is inversely mapped to a hypergraph to obtain an enhanced view.
[0031] Further, in step S5, the joint loss function is:
[0032] (27)
[0033] wherein is a hyperparameter for weighting different loss components, and satisfies ;
[0034] is a supervised binary cross-entropy loss, used to minimize the deviation of the predicted probability of the triplet association from the true label;
[0035] is a contrast consistency loss, which constrains the similarity between the node-graph level representation pair of the original hypergraph and the corresponding representation pair of the enhanced view, and strengthens the cross-view feature consistency;
[0036] is a generative reconstruction loss, which is based on the variational evidence lower bound and includes a bipartite graph association reconstruction loss and a KL divergence regularization term of the latent variable distribution and the standard normal prior.
[0037] Further, in step S6, the candidate miRNA-drug-side effect triplet is derived from the triplets in the Cartesian product of the full set of miRNAs, drugs and side effects that are not included in the benchmark data set in step S1; the prediction score is obtained by inputting the joint representation of the candidate triplet into the trained multilayer perceptron and passing through the Sigmoid activation, which is used to quantify the possibility of the association of the triplet.
[0038] Compared with the prior art, the beneficial effects of the present application are:
[0039] The application first constructs a high-quality benchmark dataset containing experimentally verified miRNA-drug-side effect (MDS) triplets. By integrating multi-source public biomedical database data, this study extracts miRNA-drug association (MDA) records from ncRNADrug V1.0 and selects 15822 experimentally verified MDA records. Drug-side effect association records are collected from Drugbank V6.0 and SIDER V4.1, and 133750 experimentally verified drug-side effect association records are selected. miRNA sequences are downloaded from miRBase V22.1, and drug SMILES data are collected from DrugBank. The data obtained above are preprocessed and filtered, and miRNA, drug and side effect information is integrated by aligning drug identifiers from different data sources to construct MDS triplets. A total of 180442 MDS triplets are obtained, covering 1578 miRNAs, 87 drugs and 5599 side effects. The study focuses on inflammation-related side effects, and the side effects are limited to this category. After filtering, 405 inflammation-related side effects are obtained. CD-HIT tool is used to cluster miRNAs with high sequence similarity, and finally 328 representative miRNAs are selected. Finally, a high-quality MDS benchmark dataset containing 37881 experimentally verified MDS triplets, covering 328 miRNAs, 87 drugs and 405 inflammation-related side effects, is constructed.
[0040] To effectively identify the high-order topological dependency relationship among miRNA, drug and side effect, the application proposes a generative augmented view strategy. The strategy uses variational hypergraph autoencoder (VHGAE) to realize parameterized generation. First, the original MDS hypergraph is losslessly mapped to an equivalent bipartite graph of "node-hyperedge", where the vertex set consists of two types of entity nodes and hyperedges, and the edge set represents the membership relationship of entity nodes to corresponding hyperedges. Second, a two-way hypergraph neural network (HyperGNN) is used as an encoder to model the entity nodes and hyperedges, learn their posterior distribution and obtain the latent representation vectors. Then, the decoder describes the association probability of entity nodes and hyperedges on the bipartite graph, and combines Gumbel-Softmax differentiable sampling to generate a new bipartite graph edge set. Finally, the sampled bipartite graph is inversely mapped to a hypergraph to obtain a generative augmented view. This generation process can mine the high-order semantics of unlabeled data under limited labeling conditions, generate augmented views according to MDS association logic and preserve key high-order structures, thereby providing reliable structure disturbance support for contrastive learning in sparse data scenarios.
[0041] To improve the accuracy of MDS association prediction, this invention proposes a joint optimization strategy that introduces supervised loss, contrastive loss, and generator loss. First, the supervised loss employs binary cross-entropy to minimize the deviation between the predicted probability and the true label, enhancing the accuracy of triple association discrimination. Second, the contrastive loss adopts a depth graph information maximization (DGI) style, using the original hypergraph's node-graph pairings as positive examples and the enhanced view's pairings as contrast examples, increasing the similarity of the former and suppressing the latter, thereby strengthening cross-view similarity. Figure One Consistency and discriminative representation. The generator loss, based on the variational evidence lower bound, improves the topological and semantic quality of the generative augmented view by maximizing the reconstruction probabilities of vertices and hyperedges in the bipartite graph and constraining the latent distribution to approximate a standard normal prior. The above three types of losses are weighted and fused into a total loss, thereby improving the sufficiency of feature learning and the prediction performance of MDS triples under sparse annotation conditions. Attached Figure Description
[0042] Figure 1 A schematic diagram illustrating the steps of a method for predicting the association between miRNAs, drugs, and side effects based on contrastive learning.
[0043] Figure 2 This is a flowchart illustrating the method for constructing the MDS triplet benchmark dataset.
[0044] Figure 3 This is a flowchart of obtaining an enhanced view from an original view based on a variational hypergraph autoencoder (VHGAE).
[0045] Figure 4 A framework diagram of a method for predicting the association between miRNAs, drugs, and side effects based on contrastive learning. Detailed Implementation
[0046] The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which are used to illustrate the principles of the present invention and are not intended to limit the scope of the present invention.
[0047] See Figure 1 This invention provides a miRNA-drug-side effect triple association prediction based on variational hypergraph encoding and contrastive learning, specifically including the following steps:
[0048] Step S1: Integrate multi-source data to construct a benchmark dataset for MDS triple identification. The specific process is as follows: Figure 2 As shown, it specifically includes:
[0049] Step S11: Download miRNA-drug association (MDA) data from the ncRNA Drug V1.0 database (https: / / www.ncrnadrug.org / ). Filter the downloaded MDA data, retain the experimentally verified association records, and eliminate the unverified predicted data. Finally, 15,822 experimentally verified miRNA-drug associations are obtained, covering 2,248 miRNAs and 183 drugs. Ensure the reliability of the association data and provide high-quality miRNA-drug interaction basis for subsequent MDS triplets construction.
[0050] Step S12: Download drug-side effect association data from the DrugBank V6.0 database (https: / / go.drugbank.com / ) and the SIDER V4.1 database (https: / / sideeffects.embl.de / ), respectively. Integrate the drug-side effect association data from the two databases, filter and retain the experimentally verified association records, and eliminate the unverified predicted data. Finally, 133,750 experimentally verified drug-side effect associations are obtained, covering 1,020 drugs and 5,599 side effects. Integrate multi-source drug-side effect information, improve data coverage and accuracy, and provide the basis for drug-side effect association in MDS triplets.
[0051] Step S13: Download miRNA sequence data from the miRBase V22.1 database (https: / / www.mirbase.org / ). Download the SMILES string of the drug from the DrugBank V6.0 database. SMILES clearly describes the molecular structure through ASCII string, providing a basis for drug molecular feature extraction. SMILES, as a linear representation of molecular structure encoded in ASCII string, facilitates unified parsing and standardized processing, and provides a basis for subsequent drug molecular feature extraction and molecular graph construction.
[0052] Step S14: Align and integrate the miRNA-drug association data obtained in step S11 and the drug-side effect association data obtained in step S12 based on the unique identifier of the drug, and then construct the initial MDS triplets. Specifically, for any miRNA, drug, and side effect, if there are experimentally verified miRNA-drug association and drug-side effect association, the miRNA-drug-side effect is considered as an initial MDS triplet, and a positive class label 1 is assigned, indicating the association between the three. After alignment and integration, the initial dataset contains 180,442 triplets, covering 1,578 miRNAs, 87 drugs, and 5,599 side effects. By linking the two types of binary relationships through the drug identifier, a preliminary association network of the three is established, laying a foundation for subsequent precise screening and modeling.
[0053] Step S15: Further filtering and optimization of the initial MDS triplets dataset, focusing on cancer-related MDS associations, while inflammation-related adverse reactions are closely related to cancer progression and treatment, thus limiting side effects to inflammation-related adverse reactions, and finally screening 405 side effects. Cluster the miRNA sequences using the CD-HIT clustering algorithm to retain representative miRNA samples, and finally obtain 328 miRNAs. Remove duplicates from the screened triplets, and retain only unique records verified by experiments, and finally obtain a benchmark dataset containing 37881 experimentally verified MDS triplets, covering 328 miRNAs, 87 drugs and 405 inflammation-related side effects.
[0054] Step S2: Based on the MSD association dataset in step S1, model the MSD association as a hypergraph structure.
[0055] Step S21: Define the node set and hyperedge set of the hypergraph to clearly define the basic components of the hypergraph.
[0056] Define hypergraph:
[0057] (1)
[0058] miRNA set
[0059] Drug set
[0060] Side effect set
[0061] Node set Composed of three types of entities: miRNA, drug, and side effect, that is, . The miRNA set obtained after screening in step S1, where represents the th miRNA, is the drug set, where represents the th drug, is the side effect set, where represents the th side effect, and each node corresponds to a specific entity.
[0062] Hyperedge set Derived from the experimentally verified MDS triplets summarized in step S1. Let the Cartesian product represent the complete set of all possible MDS triplets, then . Any hyperedge is associated with a triplet Corresponding, for depicting the synergistic relationship between miRNA, drug and side effect. The above definition provides a formal basis for subsequent hypergraph modeling and message passing.
[0063] Step S22: Construct the association matrix of the hypergraph, and quantify the membership relationship between nodes and hyperedges.
[0064] To quantify the membership relationship between nodes and hyperedges, the association matrix of the hypergraph is established . Let the first row correspond to the node , the first column correspond to the hyperedge , and the matrix element is defined as
[0065] (2)
[0066] That is, when the node belongs to the hyperedge , ; otherwise, . The association matrix records the hyperedges participated by the nodes, and provides a structural basis for the message passing of the hypergraph convolution network (HGCN), ensuring that the nodes can aggregate the feature information of the nodes in the same domain through their own hyperedges.
[0067] Step S3: Based on the hypergraph structure constructed in step S2, learn the MSD triplet representation using the hypergraph convolution network.
[0068] Step S31: To quantify the correlation between miRNAs, calculate the similarity between any two miRNAs and using cosine similarity:
[0069] (3)
[0070] Calculate for all 328 miRNAs, get a symmetric similarity matrix . Then input into the fully connected network (FCN), and perform nonlinear mapping through the ReLU activation function, output the low-dimensional feature representation matrix of miRNA , where is the number of miRNAs, is the embedding dimension.
[0071] Step S32: Based on the SMILES string of each drug, use DeepChem tools to convert SMILES into a molecular graph , where is the attribute matrix of the atoms in the molecule, Let be the adjacency matrix of chemical bonds. An isomorphic graph network (GIN) is applied to each molecular diagram to encode atomic-level characterizations, where the first is the adjacency matrix of chemical bonds. The GIN of the layer is updated as follows:
[0072] (4)
[0073] in It is the identity matrix. For a fixed scalar, and , For the first A multilayer perceptron is used. Global max pooling (GMP) is then performed on the node embeddings of the molecular graph to obtain the vector representation of a single drug, ultimately yielding the drug's feature matrix. ,in For the quantity of drugs, For the final embedding dimension.
[0074] Step S33: To quantify the correlation between side effects, the Jaccard coefficient is used to calculate the correlation between any two side effects. and Similarity:
[0075] (5)
[0076] in Indications and side effects A set of related drugs. A symmetric similarity matrix was obtained by calculating the similarity for each of the 405 side effects. .Will The input is a fully connected network (FCN), which undergoes a nonlinear mapping using the ReLU activation function, and outputs a low-dimensional feature representation matrix with side effects. ,in For the number of side effects, For the embedded dimension.
[0077] Step S34: Concatenate the feature vectors of miRNA, drug, and side effects according to entity type to form the attribute matrix of the hypergraph nodes. The formula is:
[0078] (6)
[0079] in, , , The feature matrices for miRNA, drug, and side effects are respectively concatenated vertically after being linearly projected along the same dimension, providing input for hypergraph convolution.
[0080] Step S35: The hypergraph convolutional layer achieves high-order information interaction between nodes through hyperedges. Layer calculation, first node degree and hyperedge degree matrix calculation, node degree matrix is a diagonal matrix, , hyperedge degree matrix is a diagonal matrix, , is also a diagonal matrix that stores the weight of the hyperedge. Then perform the hypergraph convolution operation, the first layer node embedding update formula is:
[0081] (7)
[0082] Nodes pass information through hyperedges, is the learnable weight matrix of the first layer, is the ReLU activation function, which enhances the model's nonlinear expression ability.
[0083] Step S36: optimize node representation by multi-layer HGCN iteration, from the initial embedding , through 3-layer HGCN layer-by-layer iteration, the first layer outputs , fusion of node itself and direct associated hyperedge information, the second layer outputs , capture the second-order high-order association of shared hyperedge nodes, the third layer outputs , integrate the global hypergraph structure to generate the final node representation . For any MDS triple , it is represented as the concatenation of the corresponding node embedding, that is
[0084] (8)
[0085] Where are the node embeddings of miRNA, drug, and side effect in the output layer of HGCN, respectively, to generate the triple representation.
[0086] Step S37: after the high-order message passing of the three-layer HGCN, the node embedding matrix is obtained, and the representation of any node is denoted as . For any triple , construct its joint representation:
[0087] (9)
[0088] And perform global average pooling on to get the graph-level vector . Then, take the "node-graph level" pair in the original view as the positive sample, and use it together with the corresponding pair in the augmented view for subsequent DGI-style contrastive learning; At the same time, the As input to the downstream discriminator, for MDS triplet correlation prediction.
[0089] Step S4: Based on the variational hypergraph autoencoder (VHGAE), an enhanced view is constructed to maintain and supplement the high-order topological semantics, and the structure and processing flow are as shown in Figure 3
[0090] Step S41: The original hypergraph is converted into a "node-hyperedge" equivalent bipartite graph, denoted as , where the bipartite graph vertex set , is the entity vertex set of the original hypergraph, is the edge set of the original hypergraph. The edge set of the bipartite graph , that is, if the entity vertex in the original hypergraph belongs to the hyperedge , then an edge connecting and is added in the bipartite graph. This conversion ensures that the high-order correlation information of the original hypergraph is completely preserved in the bipartite graph, providing a structured input for subsequent probabilistic modeling.
[0091] Step S42: Learn the hypergraph distribution using the VHGAE encoder, adopt a double-parallel hypergraph neural network (HyperGNN) as the encoder, and learn the variational distribution of the entity vertices and hyperedge vertices in the bipartite graph respectively, and output the latent vector.
[0092] For each entity vertex in the graph , the hypergraph neural network (HyperGNN) encoder outputs its variational posterior parameters, the mean vector and the logarithmic standard deviation vector , and the standard deviation is recovered by , thereby constructing the variational distribution of vertex :
[0093] (10)
[0094] where is the latent representation vector of entity vertex , is the latent dimension, is the encoder parameter.
[0095] For each hyperedge in the graph , the hypergraph neural network (HyperGNN) encoder outputs its variational posterior parameters, the mean vector and the logarithmic standard deviation vector and by reconstructing the super-edge 's variance distribution:
[0096] (11)
[0097] where is the latent vector of super-edge , is the latent dimension, and
[0098] Finally, all entity vertices and all super-edge representations are assembled into matrices.
[0099] (12)
[0100] (13)
[0101] Step S43: After obtaining the entity vertex latent variable matrix and the super-edge vertex latent variable matrix , the decoder reconstructs the "vertex-super-edge" dependency relationship on the "node-super-edge" equivalent bipartite graph , and obtains the generative augmented view through differentiable sampling.
[0102] For each pair of "entity vertex-super-edge vertex", the sigmoid function gives the dependency probability
[0103] (14)
[0104] where is the inner product of two vectors.
[0105] The log probability is mapped to the association probability in the [0, 1] interval through the sigmoid function, and the overall reconstruction probability distribution of the bipartite graph is constructed:
[0106] (15)
[0107] This distribution ensures the semantic consistency of the generated bipartite graph edge set with the original edge set.
[0108] Step S44: To solve the non-differentiable problem of discrete edge set sampling, the Gumbel-Softmax technique is used to sample the association probability of step S43 to generate a new bipartite graph edge set. For each pair, Gumbel noise , is generated, which is subject to a uniform distribution.The temperature coefficient, combined with the noise level, is used to calculate the smoothed association weights based on the log-probability of the association, using the following formula:
[0109] (16)
[0110] Edge set determination rules, if If the value exceeds the set threshold, the bipartite graph edges are retained. Otherwise, delete, and finally obtain the sampled bipartite graph. ,in This is the sampled edge set.
[0111] Step S45: Following the inverse rule of step S41, process the sampled bipartite graph... Reverting to a hypergraph structure yields a new generative enhanced view. The hypergraph vertex set maintains the same shape as the original graph. Figure One To, that is Hypergraph Hyperedgeset That is, for each hyperedge vertex in a bipartite graph Connect all the entity vertices to form a new hyperedge. The above steps generate It can preserve the high-order association semantics of the original hypergraph and generate a comparative view.
[0112] Step S5: Based on the contrastive learning framework, a joint optimized supervised loss is set for model training to improve the model's discriminative ability and generalization performance. The overall training process and the relationship between each functional module of this invention are as follows: Figure 4 As shown.
[0113] Step S51: Perform the HGCN encoding process of step S3 on the enhanced view generated in step S4 to obtain the node embedding of the enhanced view. The matrix yields the enhanced node embeddings. Embedding of original hypergraph nodes Perform global average pooling to generate a graph-level representation. ,Will As a positive example ( For nodes in the original view Embedded). As a negative example ( To enhance nodes in the view Embedding), constructing contrastive learning sample pairs. Define contrastive loss. :
[0114] (17)
[0115] in, This represents the total number of nodes in the hypergraph. is a discriminator composed of a bilinear function and a sigmoid activation, which is used to calculate the similarity between node-level representation and graph-level representation. This loss makes the model learn more discriminative features by penalizing high similarity of negative sample pairs and rewarding high similarity of positive sample pairs.
[0116] Step S52: For MDS correlation prediction, we utilize miRNA , drug and side effect learning embeddings to output their correlation probabilities via a scoring function:
[0117] (18)
[0118] The binary cross-entropy loss is used as the supervised task loss , which is calculated as:
[0119] (19)
[0120] where is the training set of triplets, is the true label of the triplet , and is the correlation probability predicted by the model.
[0121] Step S53: The reconstruction loss term aims to let the decoder restore the “node-hyperedge relationship” in the original hypergraph based on the latent variables, so as to measure the structural similarity between the generated view and the original view. The encoder outputs the latent variable distribution of the node and the hyperedge respectively.
[0122] 、 The decoder gives the probability that the node belongs to the hyperedge . Accordingly, the reconstruction loss is defined as the negative expectation of the log-likelihood on the posterior distribution, that is:
[0123] (20)
[0124] To suppress overfitting and preserve the necessary randomness for the generated view, a KL divergence regularization term is added to constrain each latent distribution to a standard Gaussian prior . Specifically, the KL divergence of the latent posterior of a single node and the prior can be written as:
[0125] (21)
[0126] The vertex term is obtained by averaging over all vertices:
[0127] (twenty two)
[0128] Similarly, for a single hyperedge Latent variables posterior with prior The KL divergence is:
[0129] (twenty three)
[0130] The superedge term is obtained by averaging over all superedges:
[0131] (twenty four)
[0132] The above two items are combined into KL divergence regularization, which is used to constrain the latent variables of vertices and hyperedges to approximate the standard prior, thereby improving the generalization ability of the model and stabilizing the quality of generated views.
[0133] Step S54: The overall form of the generation loss can be expressed as follows: the lower bound of evidence (ELBO) is composed of the expectation of the log-likelihood on the posterior distribution minus the KL divergence terms for the vertices and hyperedges, i.e.
[0134] (25)
[0135] (26)
[0136] Step S55: To simultaneously consider the performance of supervised discrimination and cross-view... Figure One To improve consistency and structural quality of generative augmented views, this invention combines three terms: supervised binary cross-entropy loss, DGI style contrast consistency loss, and the generator's evidence lower bound (ELBO) term. The training objective is defined as minimizing the overall loss.
[0137] (27)
[0138] in It is a hyperparameter that weighs different loss components and satisfies .
[0139] Step S6: Obtain the MSD triplet prediction score
[0140] Step S61: Combine the validation phase samples with the unknown combinations to form a candidate set. From the complete set of triplet After removing the labeled positive examples from the training set, the resulting unknown combinations are merged with the validation samples. It also performs deduplication and label verification to ensure that the candidate set and the training positive examples are strictly disjoint.
[0141] Step S62: Call the trained HGCN encoder on the original view to get the node embedding matrix . For any candidate triple , construct the joint representation according to the rule in step S38.
[0142] Step S63: Input the triple joint representation into the trained discriminator, a multi-layer perceptron with a Sigmoid output layer, to get the association probability in the interval , which is taken as the predicted score of the triple.
[0143] Step S64: Select a threshold on the validation set to determine whether there is an association between the triples . Specifically, triples with predicted scores higher than the set threshold are labeled as associated, while those with lower scores are labeled as unassociated. This threshold setting method can quantitatively determine and confirm the association between each group of miRNA, drug, and side effect, providing clear guidance and foundation for further research.
[0144] The above description is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto, any changes or replacements within the technical range disclosed by the present application can be easily thought of by those skilled in the art, which should be covered within the protection scope of the present application.
Claims
1. A method for predicting miRNA-drug-side effect ternary associations, characterized in that, Includes the following steps: S1: Integrate multi-source data to construct a benchmark dataset for the identification of miRNA-drug-side effect triples; S2: Based on the benchmark dataset in step S1, model the miRNA-drug-side effect association as a hypergraph structure; S3: Based on the hypergraph structure constructed in step S2, learn the miRNA-drug-side effect triplet representation using a hypergraph convolutional network; S4: Constructing an enhanced view based on a variational hypergraph autoencoder; S5: Based on the contrastive learning framework, the hypergraph structure and augmented view are treated as a whole, and end-to-end training is performed using a joint loss function consisting of supervised binary cross-entropy loss, contrastive consistency loss and generative reconstruction loss to improve the model’s discriminative ability and generalization performance. S6: Use the trained model to predict the association between candidate miRNA-drug-side effect triples and output the prediction score.
2. The miRNA-drug-side effect ternary association prediction method according to claim 1, characterized in that, In step S1, the integration of multi-source data includes: extracting experimentally validated miRNA-drug association data from the ncRNADrug database, collecting experimentally validated drug-side effect association data from DrugBank and SIDER databases, downloading miRNA sequence data from miRBase, and obtaining drug SMILES string data from DrugBank.
3. The method according to claim 2, characterized in that, The construction of the benchmark dataset also includes: aligning and integrating miRNA-drug association data and drug-side effect association data using drug unique identifiers to construct an initial MDS triple dataset; further screening and optimization of the initial MDS triple dataset to finally obtain the benchmark dataset of MDS triples.
4. The miRNA-drug-side effect ternary association prediction method according to claim 1, characterized in that, In step S2, the hypergraph structure is defined as , where: node set From miRNA set Drug collection Side effects collection Composition, that is Hyperedge set Each hyperedge in This corresponds to a miRNA-drug-side effect triplet in the baseline dataset in step S1. This is used to characterize the synergistic relationship among the three.
5. The miRNA-drug-side effect ternary association prediction method according to claim 4, characterized in that, In step S2, the hypergraph structure further includes an association matrix. , where matrix elements Represents a node Belongs to superedge , Represents a node Not a hyperedge This is used to quantify the dependency relationship between nodes and hyperedges.
6. The miRNA-drug-side effect ternary association prediction method according to claim 1, characterized in that, In step S3, learning the miRNA-drug-side effect triplet representation using a hypergraph convolutional network specifically includes: (1) Extract the feature representations of miRNA, drug and side effects respectively: The similarity between miRNAs is calculated by cosine similarity and the miRNA features are obtained by mapping through a fully connected network. The molecular map is constructed based on drug SMILES and the drug features are obtained by encoding through an isomorphic graph network. The similarity between side effects is calculated by Jaccard coefficient and the side effect features are obtained by mapping through a fully connected network. (2) The above features are concatenated into a hypergraph node attribute matrix and input into the hypergraph convolutional network. The higher-order association features of the nodes are learned through multi-layer hypergraph convolutional operations. (3) The node features of miRNA, drug and side effects are spliced together to obtain the joint representation of miRNA-drug-side effects triple.
7. The miRNA-drug-side effect ternary association prediction method according to claim 1, characterized in that, In step S4, constructing the enhanced view based on the variational hypergraph autoencoder specifically includes: (1) Transform the hypergraph constructed in step S2 into a node-hyperedge equivalent bipartite graph, where the vertex set contains entity nodes and hyperedges, and the edge set represents the membership relationship between entity nodes and hyperedges; (2) A dual-path hypergraph neural network is used as the encoder to learn the latent variable distributions of entity nodes and hyperedges respectively; (3) The association probability between entity nodes and hyperedges is reconstructed by the decoder, and a new bipartite graph edge set is generated by combining Gumbel-Softmax differentiable sampling; (4) Inverse map the new bipartite graph to a hypergraph to obtain an enhanced view.
8. The miRNA-drug-side effect ternary association prediction method according to claim 1, characterized in that, In step S5, the joint loss function is: (27) in It is a hyperparameter that weighs different loss components and satisfies ; To supervise the binary cross-entropy loss, it is used to minimize the deviation between the predicted probability of triple association and the true label; To compare consistency loss, cross-view feature consistency is enhanced by constraining the similarity between node-graph level representation pairs of the original hypergraph and corresponding representation pairs of the enhanced view. To generate the reconstruction loss, based on the variational evidence lower bound, it includes the bipartite graph association reconstruction loss and the KL divergence regularization term of the latent variable distribution and the standard normal prior.
9. The method according to claim 1, characterized in that, In step S6, the candidate miRNA-drug-side effect triplet is derived from the Cartesian product of the complete set of miRNA, drug, and side effects that is not included in the benchmark dataset in step S1; the predicted score is obtained by inputting the joint representation of the candidate triplet into a trained multilayer perceptron and activating it with Sigmoid, and is used to quantify the probability of triplet association.