Drug-lncRNA relation prediction method and system based on embedding constraint
By constructing a lncRNA-drug bipartite graph and introducing the LDA-GNN and LDA-DHG models with embedded constraint strategies, the problem of insufficient capture of complex graph structures by existing methods in predicting the association between long non-coding RNA and drugs is solved, achieving higher prediction accuracy and biological interpretability.
Patent Information
- Application Number
- CN202510701935.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-05-28
AI Technical Summary
Existing methods have difficulty in fully capturing complex graph structure relationships when predicting the association between long non-coding RNA and drugs, resulting in insufficient prediction performance and discrimination capabilities.
A drug-lncRNA relationship prediction method based on embedding constraints was adopted. By constructing a lncRNA-drug bipartite graph, obtaining the similarity matrix and introducing the LDA-GNN and LDA-DHG models with embedding constraint strategies, the proximity of related nodes in the embedding space was ensured, thereby improving the prediction performance.
It significantly improves the model's predictive performance and discriminative ability, improves the quality of learned representations, enhances the biological interpretability of prediction results, and supports deeper molecular pathology analysis and target identification.
Smart Images

Figure CN120636602A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of drug-lncRNA relationship prediction, and particularly relates to a drug-lncRNA relationship prediction method and system based on embedded constraints. Background Art
[0002] Advances in high-throughput experimental techniques and computational methods have greatly advanced our ability to understand drug mechanisms of action, particularly by revealing the complex interactions between drugs and non-coding RNAs at the molecular level. In recent years, predicting associations between long non-coding RNAs (lncRNAs) and drugs (lncRNA-drug associations, or LDRAs) has become a key research direction in drug repurposing and personalized therapy. LncRNAs are a class of RNA molecules that do not encode proteins and are widely involved in regulating gene expression, cellular function, and disease mechanisms. Studies have shown that lncRNAs can significantly influence the pharmacological properties of drugs by affecting the expression of drug targets or regulating key signaling pathways.
[0003] Currently, several databases systematically integrate experimentally validated LDRAs. For example, D-lnc and NcRNADrug include data on the effects of drugs on lncRNA expression and lncRNAs associated with drug resistance. These data provide a valuable resource for in-depth study of the functional relationships between lncRNAs and drugs. However, experimental methods are often costly and time-consuming, limiting their application in large-scale prediction tasks. Therefore, the development of computational LDRAs prediction models has become an urgent issue.
[0004] Various machine learning methods have been applied to LDRAs prediction. For example, Wang et al. used the ElasticNet (EN) regression model to integrate genomic data from various cancer cell lines with drug response data to identify LDRAs. Ha et al. proposed a matrix decomposition method based on lncRNA expression profiling (EMFLDA) to integrate heterogeneous biological data. Chen et al. proposed a random walk (RWR)-based LDRAs prediction method that takes into account the importance of local structure. Although these methods have achieved some success, they often have certain limitations in capturing high-dimensional and nonlinear features and integrating various biological data.
[0005] In recent years, deep learning (DL) technology has gained prominence due to its ability to learn hierarchical representations and capture complex patterns in data. Tang et al. proposed a GCN framework to integrate multiple similarities between lncRNAs and drugs and construct a unified graph for end-to-end prediction. Liu et al. proposed a method that combines a collaborative learning framework and graph structure learning to model the synergistic change pattern between lncRNAs and drugs. Ha et al. used a matrix decomposition method based on lncRNA expression profile data and drug response to fuse multiple heterogeneous biological features. Although existing deep learning methods have improved the prediction performance to a certain extent, they still cannot fully explore the potential dependency characteristics between molecular entities because they rely on static feature representations and find it difficult to fully capture the complex graph structure relationship between lncRNAs and drugs.
[0006] Among deep learning techniques, graph neural networks (GNNs) are becoming increasingly popular in molecular prediction tasks because they can effectively model molecular entities as nodes and capture their relationships as edges in a graph structure. For example, Xu et al. introduced a LDRAs recognition model (LDA-GNN) based on a graph attention network (GAT) and a graph convolutional network (GCN). Xu et al. developed a model (LDA-DHG) for identifying DTIs, which constructed a dynamic heterogeneous graph based on drug-drug and target-target similarity networks and proposed a new progressive learning strategy to improve the accuracy of prediction. The model was also applied to the field of LDRA prediction. Yao et al. proposed a prediction model by integrating GCN and Transformer, which constructed a graph adjacency matrix based on intra-class similarity and inter-class connections between lncRNA, miNRA and diseases.
[0007] However, traditional GNN-based methods still have limitations in fully capturing the spatial constraints between related nodes. Specifically, they do not always enforce the proximity of biologically relevant entities, which may hinder the quality of learned representations. How to overcome this limitation is a technical problem that needs to be solved urgently. Summary of the Invention
[0008] To address the problems existing in the prior art, the present invention provides a drug-lncRNA relationship prediction method and system based on embedded constraints, aiming to address the shortcomings of existing methods in exploring the geometric relationships between molecular entities and improve the prediction performance and discrimination ability of the model.
[0009] To achieve the above object, the present invention provides the following solutions:
[0010] A method for predicting drug-lncRNA relationships based on embedded constraints, the method comprising:
[0011] S1. Collect and preprocess lncRNA-drug association data to construct a dataset;
[0012] S2. Based on the data set, construct a lncRNA-drug association matrix and extract a lncRNA-drug bipartite graph;
[0013] S3. Based on the lncRNA-drug association matrix, obtaining a lncRNA similarity matrix and a drug similarity matrix, and performing dimensionality reduction processing on the lncRNA similarity matrix and the drug similarity matrix;
[0014] S4. Inputting the lncRNA-drug bipartite graph, the lncRNA similarity matrix after dimensionality reduction, and the drug similarity matrix into an LDA-GNN model and an LDA-DHG model with an embedded constraint strategy to obtain a probability score of the association between the lncRNA and the drug;
[0015] S5. Based on the probability score, predict the relationship between drug and lncRNA.
[0016] Preferably, the S3 includes:
[0017] G l (l i ,l j )=exp(-r l ||Y(l i )-Y(l j )|| 2 )
[0018] G d (d i ,d j )=exp(-r d ||Y(d i )-Y(d j )|| 2 )
[0019] Among them, G l (l i ,l j ) and G d (d i ,d j ) are lncRNAl i and lncRNAl j Gaussian similarity and drug d i With drug d j Gaussian similarity, r l and r d are the scale functions of the Gaussian kernel function in controlling the similarity between lncRNA and drugs, exp(·) is the natural exponential function, and Y(l i)、Y(l j )、Y(d i ) and Y(d j ) are lncRNA l i 、lncRNAl j , drugs d i and drugs j The eigenvector representation of ||Y(l i )-Y(l j )|| 2 lncRNAl i and lncRNAl j The square of the Euclidean distance between the eigenvectors of ||Y(d i )-Y(d j )|| 2 Indicates drug d i With drug d j The squared Euclidean distance between the eigenvectors of .
[0020] Preferably, the S4 includes:
[0021]
[0022] in, represents the output features of the lth layer, represents the output feature of the l-1 layer, W (l) is the weight matrix of the lth layer, b (l) is the bias term of the lth layer, and ReLU(·) is the activation function.
[0023] Preferably, the S4 further includes: training the LDA-GNN model and the LDA-DHG model with the embedded constraint strategy:
[0024] The lncRNA-drug association data, the lncRNA similarity matrix after dimensionality reduction, and the drug similarity matrix are input as edges and nodes of the graph structure into the LDA-GNN model constructed based on GAT and GCN, and the node update method is as follows:
[0025]
[0026] in, Represents the feature vector of node j at layer 0 in the LDA-GNN model, represents the feature vector of node i in the first layer of the LDA-GNN model, d(i) represents the degree of node i, d(j) represents the degree of node j, N(i) represents the set of all neighbors of node i, and w 1 Represents the parameter matrix that GCN needs to learn;
[0027] The lncRNA-drug association data, the lncRNA similarity matrix after dimensionality reduction, and the drug similarity matrix are input as edges and nodes of the graph structure into the LDA-DHG model constructed based on GAT and GCN, and the node update method is as follows:
[0028]
[0029] in, represents the feature of node i at the lth layer in the LDA-DHG model, Represents the feature of node i in the l+1th layer in the LDA-DHG model, Agg is the aggregation operation, and MLP is the multi-layer perceptron used for feature update;
[0030] Introduce the attention mechanism and calculate the attention coefficient between nodes:
[0031]
[0032] Among them, e ij is the attention coefficient of node i to neighbor node j, is the feature representation of node i in the kth layer, is the feature representation of node j in the kth layer, a is the shared attention mechanism for calculating the attention coefficient; W (k) represents the parameter matrix that needs to be learned in the kth layer, and σ represents the activation function;
[0033] Use the Softmax function to normalize the attention coefficient:
[0034]
[0035] Among them, a ij is the normalized value of the attention coefficient, e im For node i to all its neighbor nodes m∈N i The attention coefficient;
[0036] Weighted sum:
[0037]
[0038] Preferably, the S4 further includes:
[0039]
[0040] Among them, y i represents the true label of the i-th sample, y i ' represents the predicted value of the i-th sample, x a represents the embedding vector of the a-th associated pair, y brepresents the embedding vector of the b-th associated pair, α and β are hyperparameters, N is the different classification results, and n is the dimension of each vector.
[0041] The present invention also provides a drug-lncRNA relationship prediction system based on embedded constraints, which is used to implement the aforementioned drug-lncRNA relationship prediction method based on embedded constraints. The system includes: an acquisition module, a bipartite graph extraction module, a similarity matrix acquisition module, a probability acquisition module and a prediction module;
[0042] The acquisition module is used to collect lncRNA-drug association data and perform preprocessing to construct a data set;
[0043] The bipartite graph extraction module is used to construct a lncRNA-drug association matrix based on the data set and extract a lncRNA-drug bipartite graph;
[0044] The similarity matrix acquisition module is used to obtain a lncRNA similarity matrix and a drug similarity matrix based on the lncRNA-drug association matrix, and perform dimensionality reduction processing on the lncRNA similarity matrix and the drug similarity matrix;
[0045] The probability acquisition module is used to input the lncRNA-drug bipartite graph and the lncRNA similarity matrix and the drug similarity matrix after dimensionality reduction into the LDA-GNN model and the LDA-DHG model with an embedded constraint strategy to obtain the probability score of the association between lncRNA and drug;
[0046] The prediction module is used to predict the relationship between drug and lncRNA based on the probability score.
[0047] Preferably, the similarity matrix acquisition module includes:
[0048] G l (l i ,l j )=exp(-r l ||Y(l i )-Y(l j )|| 2 )
[0049] G d (d i ,d j )=exp(-r d ||Y(d i )-Y(d j )|| 2 )
[0050] Among them, G l (li ,l j ) and G d (d i ,d j ) are lncRNAl i and lncRNAl j Gaussian similarity and drug d i With drug d j Gaussian similarity, r l and r d are the scale functions of the Gaussian kernel function in controlling the similarity between lncRNA and drugs, exp(·) is the natural exponential function, and Y(l i )、Y(l j )、Y(d i ) and Y(d j ) are lncRNA l i 、lncRNAl j , drugs d i and drugs j The eigenvector representation of ||Y(l i )-Y(l j )|| 2 lncRNAl i and lncRNAl j The square of the Euclidean distance between the eigenvectors of ||Y(d i )-Y(d j )|| 2 Indicates drug d i With drug d j The squared Euclidean distance between the eigenvectors of .
[0051] Preferably, the probability acquisition module includes:
[0052]
[0053] in, represents the output features of the lth layer, represents the output feature of the l-1 layer, W (l) is the weight matrix of the lth layer, b (l) is the bias term of the lth layer, and ReLU(·) is the activation function.
[0054] Preferably, the probability acquisition module further comprises: training an LDA-GNN model and an LDA-DHG model with an embedded constraint strategy:
[0055] The lncRNA-drug association data, the lncRNA similarity matrix after dimensionality reduction, and the drug similarity matrix are input as edges and nodes of the graph structure into the LDA-GNN model constructed based on GAT and GCN, and the node update method is as follows:
[0056]
[0057] in, Represents the feature vector of node j at layer 0 in the LDA-GNN model, represents the feature vector of node i in the first layer of the LDA-GNN model, d(i) represents the degree of node i, d(j) represents the degree of node j, N(i) represents the set of all neighbors of node i, and w 1 Represents the parameter matrix that GCN needs to learn;
[0058] The lncRNA-drug association data, the lncRNA similarity matrix after dimensionality reduction, and the drug similarity matrix are input as edges and nodes of the graph structure into the LDA-DHG model constructed based on GAT and GCN, and the node update method is as follows:
[0059]
[0060] in, represents the feature of node i at the lth layer in the LDA-DHG model, Represents the feature of node i in the l+1th layer in the LDA-DHG model, Agg is the aggregation operation, and MLP is the multi-layer perceptron used for feature update;
[0061] Introduce the attention mechanism and calculate the attention coefficient between nodes:
[0062]
[0063] Among them, e ij is the attention coefficient of node i to neighbor node j, is the feature representation of node i in the kth layer, is the feature representation of node j in the kth layer, a is the shared attention mechanism for calculating the attention coefficient; W (k) represents the parameter matrix that needs to be learned in the kth layer, and σ represents the activation function;
[0064] Use the Softmax function to normalize the attention coefficient:
[0065]
[0066] Among them, a ij is the normalized value of the attention coefficient, e imFor node i to all its neighbor nodes m∈N i The attention coefficient;
[0067] Weighted sum:
[0068]
[0069] Preferably, the probability acquisition module further includes:
[0070]
[0071] Among them, y i represents the true label of the i-th sample, y i ' represents the predicted value of the i-th sample, x a represents the embedding vector of the a-th associated pair, y b represents the embedding vector of the b-th associated pair, α and β are hyperparameters, N is the different classification results, and n is the dimension of each vector.
[0072] Compared with the prior art, the present invention has the following beneficial effects:
[0073] This paper introduces an embedding constraint strategy to address the shortcomings of existing methods in exploring the geometric relationships between molecular entities. By ensuring the proximity of related nodes in the embedding space, the model's predictive performance and discriminative capabilities are significantly improved, enhancing the quality of learned representations. By introducing the embedding constraint strategy, the present invention achieved significant performance improvements in two graph neural network models, LDA-GNN and LDA-DHG, demonstrating higher AUC, AUPR, and Pre@k metrics in the LDRA prediction task. This distance constraint ensures that confirmed molecular entities such as lncRNAs, drugs, and targets are positioned closer in the feature space, resulting in more accurate predictions.
[0074] Experimental results demonstrate that the embedding constraint strategy exhibits robustness across diverse data scales, is adaptable to both small and large datasets, and has strong potential for widespread application. By optimizing the embedding space, this method helps accurately distinguish between positive and negative samples, thereby improving the biological interpretability of prediction results. This supports deeper molecular pathology analysis and target identification, contributing to more effective drug discovery and disease research. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0076] Figure 1This is a flow chart of a method for predicting drug-lncRNA relationships based on embedded constraints according to an embodiment of the present invention;
[0077] Figure 2 This is a line graph of the two models in the LDRAs prediction task after the embedding constraint strategy is introduced in the embodiment of the present invention; (a) and (b) are line graphs of the AUC and AUPR values of the two models on datasets 1, 2, and 3, respectively;
[0078] Figure 3 Schematic diagram of sample distance distribution in LDRAs prediction task after the two models in the embodiment of the present invention introduce embedded constraints;
[0079] Figure 4 This is a schematic diagram of the embedding similarity of association pairs of two models in the LDRAs prediction task after the embedding constraint is introduced in an embodiment of the present invention. DETAILED DESCRIPTION
[0080] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0081] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0082] Example 1
[0083] like Figure 1 As shown, the present invention provides a drug-lncRNA relationship prediction method based on embedded constraints, comprising:
[0084] S1. Collect and preprocess lncRNA-drug association data to construct a dataset;
[0085] S2. Based on the dataset, construct the lncRNA-drug association matrix and extract the lncRNA-drug bipartite graph;
[0086] S3. Based on the lncRNA-drug association matrix, obtain the lncRNA similarity matrix and the drug similarity matrix, and perform dimensionality reduction processing on the lncRNA similarity matrix and the drug similarity matrix;
[0087] S4. Input the lncRNA-drug bipartite graph and the lncRNA similarity matrix and drug similarity matrix after dimensionality reduction into the LDA-GNN model and LDA-DHG model with an embedded constraint strategy to obtain the probability score of the association between lncRNA and drug;
[0088] S5. Based on the probability score, drug-lncRNA relationship prediction is achieved.
[0089] The following is an explanation based on the specific LncRNA-drug association prediction task:
[0090] Input and output data description:
[0091] The lncRNA-drug association dataset was downloaded from http: / / www.jianglab.cn / D-lnc / index.jsp and duplicate associations were removed. The Validated dataset contains 4691 lncRNAs, 48 drugs, and 4791 associations, the cMap dataset contains 129 lncRNAs, 1279 drugs, and 15804 associations, and the GEO dataset contains 2360 lncRNAs, 115 drugs, and 28487 associations.
[0092] (1) L = {l1,l2,...,l m} represents the lncRNA dataset, D={d1,d2,...,d n} represents the drug dataset, m represents the number of LncRNAs in the dataset L, and n represents the number of drugs in the dataset D;
[0093] (2) Construct an association matrix to represent the association information between LncRNA and drug, Y ij =0 indicates lncRNA l i and drugs j Unknown relationship, Y ij =1 indicates lncRNAl i and drugs j There is a correlation. The known correlation is regarded as a positive sample, and the same number of unknown correlations as the positive samples is regarded as a negative sample.
[0094] (3) According to the lncRNA-drug association matrix G = (U, V, E), the lncRNA-drug bipartite graph is extracted, where U = {l1,l2,....,l m} is the lncRNA node set, V={d1,d2,....,d n} is the drug node set, and the edge set E consists of all node pairs satisfying A(i,j)=1, indicating that lncRNAl i and drugs jThere is an undirected edge between them;
[0095] (4) Based on the lncRNA-drug association matrix, the Gaussian interaction feature kernel function, which is widely used to calculate vector similarity, is used to calculate the lncRNA similarity matrix and the drug similarity matrix, respectively.
[0096] The formula for calculating lncRNA similarity is:
[0097] G l (l i ,l j )=exp(-r l ||Y(l i )-Y(l j )|| 2 )
[0098] The formula for calculating drug similarity is:
[0099] G d (d i ,d j )=exp(-r d ||Y(d i )-Y(d j )|| 2 )
[0100] Among them, G l (l i ,l j ) and G d (d i ,d j ) are lncRNAl i and lncRNAl j Gaussian similarity and drug d i With drug d j Gaussian similarity, r l and r d are the scale functions of the Gaussian kernel function in controlling the similarity between lncRNA and drugs, respectively, and are used to control the “smoothness” of the similarity function; exp(·) is the natural exponential function, and Y(l i )、Y(l j )、Y(d i ) and Y(d j ) are lncRNAl i 、lncRNAl j , drugs d i and drugs j The eigenvector representation of ||Y(l i )-Y(l j )|| 2 lncRNAli and lncRNAl j The square of the Euclidean distance between the eigenvectors of ||Y(d i )-Y(d j )|| 2 Indicates drug d i With drug d j The squared Euclidean distance between the eigenvectors of .
[0101] For the LDA-DHG model, two methods, global sorting and local sorting, are used to construct the similarity matrix between nodes. Global sorting:
[0102]
[0103] Global sorting sorts all elements of the similarity matrix and calculates the percentile value of each element, where N is the total number of nodes, || is the indicator function, and the similarity S(i,j) is the similarity value between node i and node j, and S(i,k) is the similarity value between node i and other nodes k∈{1,N};
[0104] Local sorting:
[0105]
[0106] Local sorting sorts the elements in each row and calculates the relative rank of each element in the row, where M is the number of elements in each row of the matrix.
[0107] (5) The row vectors of the lncRNA similarity matrix are used as the eigenvectors of lncRNA, and the row vectors of the drug similarity matrix are used as the eigenvectors of drugs, and PCA dimensionality reduction is performed on them;
[0108] (6) Input the lncRNA-drug bipartite graph, the reduced lncRNA similarity matrix, and the drug similarity matrix;
[0109] (7) The lncRNA feature h obtained by neural network training i and drug characteristics h j Input into the deep neural network, the deep neural network consists of L layers, and the calculation process of each layer is:
[0110]
[0111] in, That is, the initial input is the concatenation feature of the node pair, represents the output features of the lth layer, represents the output feature of the l-1 layer, W (l) is the weight matrix of the lth layer, b (l)is the bias term of the lth layer, ReLU(·) is the activation function, and the number of layers is set to 3. After the neural network transformation, the activation function output is input to obtain the probability score of the association between lncRNA and drug.
[0112] Model training:
[0113] In order to demonstrate the generalization ability of the embedding constraint strategy, the present invention introduces the strategy into the currently commonly used model for drug-lncRNA relationship prediction. Since there are few existing methods in the field of relationship prediction, the present invention selects two models with better performance, namely the LDA-GNN model based on the graph attention network (GAT) and the graph convolutional network (GCN), and the dynamic heterogeneous graph LDA-DHG constructed based on the lncRNA-lncRNA and drug-drug similarity network. The present invention introduces the embedding constraint strategy into two existing GNN models for predicting drug-lncRNA relationships. Specifically, a regularization method is designed when the network updates the node embedding to make the embedding vectors of node pairs known to be associated closer. The following is a detailed introduction to the network architecture of the two models:
[0114] (1) The LDA-GNN model consists of a layer of graph convolutional network (GCN) and a multi-layer graph attention layer (GAT). A layer of GCN is used to aggregate the information of neighboring nodes to update its own features. The lncRNA-drug association data, the lncRNA similarity matrix after dimensionality reduction, and the drug similarity matrix are input into the LDA-GNN model as the edges and nodes of the graph structure. The node update method is:
[0115]
[0116] Among them, σ represents the activation function, Represents the feature vector of node j in the 0th layer (input layer) of the LDA-GNN model, represents the feature vector of the first layer (input layer) of node i in the LDA-GNN model, d(i) represents the degree of node i, d(j) represents the degree of node j, N(i) represents the set of all neighbors of node i, and w 1 Represents the parameter matrix that GCN needs to learn.
[0117] (2) The LDA-DHG model is based on the GAT and GCN networks, in which the features of the nodes are updated by aggregating the features of the neighboring nodes. The lncRNA-drug association data, the lncRNA similarity matrix after dimensionality reduction, and the drug similarity matrix are input into the LDA-DHG model as the edges and nodes of the graph structure. The node update method is:
[0118]
[0119] in, represents the feature of node i at the lth layer in the LDA-DHG model, represents the feature of node i at the l+1th layer in the LDA-DHG model, N(i) represents the set of all neighbors of node i, Agg is the aggregation operation, and MLP is a multi-layer perceptron used for feature update.
[0120] (3) Although GAT is similar to GCN in that both update node features by aggregating information from neighboring nodes, GAT allows different weights to be assigned to different neighboring nodes during the aggregation process and introduces an attention mechanism so that more important nodes receive more attention. Taking nodes i and j as an example, calculate the attention coefficient between them:
[0121]
[0122] Among them, e ij is the attention coefficient of node i to neighbor node j, is the feature representation of node i in the kth layer, is the feature representation of node j in the kth layer, a is the shared attention mechanism for calculating the attention coefficient, its initialization is random, and the parameters are updated through training, W (k) Represents the parameter matrix that needs to be learned in the kth layer, σ represents the activation function, and LeakyReLU is selected as the activation function.
[0123] (4) In order to compare the attention coefficients between different nodes, the Softmax function is used to normalize the attention coefficients:
[0124]
[0125] Among them, a ij is the normalized value of the attention coefficient, e ij is the attention coefficient of node i to neighbor node j, e im For node i to all its neighbor nodes m∈N i The attention coefficient is used for softmax normalization.
[0126] After calculating the attention coefficient between node i and its neighboring nodes, we can assign different weights to the neighboring nodes of node i, and then calculate the weighted sum of the neighboring nodes as the input of the next layer:
[0127]
[0128] Distance constraint strategy:
[0129] The present invention is based on the improvement of the existing distance constraint loss strategy. By utilizing the known relationship of the association pairs in the data set, the proximity between biologically related nodes in the feature space is explicitly enhanced, thereby improving the accuracy and generalization ability of drug-lncRNA relationship prediction. In the feature update process of the LDA-GNN model and the LDA-DHG model, the distance constraint strategy is introduced respectively, and the association pairs that have been proven to have a relationship are set as positive samples, and the other association pairs are set as negative samples. By reducing the distance between the embedding vectors of positive samples during training, the effective separation of positive and negative samples in the feature space is promoted. During the embedding vector update process, the feature representations of lncRNA and drug nodes with known association relationships will become more similar, thereby further optimizing the feature distribution and enhancing the biological interpretability and predictive performance of the model. The distance constraint ensures that the molecular entities of lncRNA and drugs are closer in the feature space. Specifically, in the loss function:
[0130]
[0131] Among them, y i represents the true label of the i-th sample, y i ' represents the predicted value of the i-th sample, x a represents the embedding vector of the a-th associated pair, y b Represents the embedding vector of the b-th association pair, α and β are hyperparameters, N is the different classification results, that is, whether there is an association, and n is the dimension of each vector.
[0132] Different hyperparameters will affect the prediction performance of the model. To achieve the best performance, we tried different hyperparameters and finally selected a set of hyperparameters with the best model prediction performance, which are α of 0.5, β of 0.5, dropout of 0.2, epoch of 40, batch size of 32, and learning rate of 0.001.
[0133] The effect of the present invention is further illustrated by experiments, using the prediction performance of the five-fold cross validation algorithm, the results are as follows Figure 2-Figure 4 shown.
[0134] In summary, this paper introduces an embedding constraint strategy to address the shortcomings of existing methods in exploring geometric relationships between molecular entities. By ensuring the proximity of related nodes in the embedding space, the model's predictive performance and discriminative capabilities are significantly improved. By introducing this embedding constraint strategy, the present invention achieves significant performance improvements in both the LDA-GNN and LDA-DHG graph neural network models, demonstrating higher AUC, AUPR, and Pre@k metrics in the LDRA prediction task.
[0135] Experimental results demonstrate that the embedding constraint strategy exhibits robustness across diverse data scales, is adaptable to both small and large datasets, and has strong potential for widespread application. By optimizing the embedding space, this method facilitates accurate differentiation between positive and negative samples, thereby enhancing the biological interpretability of predictions and supporting deeper molecular pathology analysis and target identification.
[0136] Example 2
[0137] The present invention also provides a drug-lncRNA relationship prediction system based on embedded constraints, which is used to implement the drug-lncRNA relationship prediction method based on embedded constraints of the above embodiment. The system includes: an acquisition module, a bipartite graph extraction module, a similarity matrix acquisition module, a probability acquisition module and a prediction module;
[0138] The acquisition module is used to collect lncRNA-drug association data and perform preprocessing to construct a dataset;
[0139] The bipartite graph extraction module is used to construct the lncRNA-drug association matrix and extract the lncRNA-drug bipartite graph based on the dataset;
[0140] A similarity matrix acquisition module is used to obtain the lncRNA similarity matrix and the drug similarity matrix based on the lncRNA-drug association matrix, and perform dimensionality reduction processing on the lncRNA similarity matrix and the drug similarity matrix;
[0141] The probability acquisition module is used to input the lncRNA-drug bipartite graph and the lncRNA similarity matrix and drug similarity matrix after dimensionality reduction into the LDA-GNN model and LDA-DHG model with embedded constraint strategy to obtain the probability score of the association between lncRNA and drug;
[0142] The prediction module is used to predict the relationship between drug and lncRNA based on probability scores.
[0143] Furthermore, the similarity matrix acquisition module includes:
[0144] G l (l i ,l j )=exp(-r l ||Y(l i )-Y(l j )|| 2 )
[0145] G d (d i ,d j )=exp(-r d ||Y(d i )-Y(dj )|| 2 )
[0146] Among them, G l (l i ,l j ) and G d (d i ,d j ) are lncRNAl i and lncRNAl j Gaussian similarity and drug d i With drug d j Gaussian similarity, r l and r d are the scale functions of the Gaussian kernel function in controlling the similarity between lncRNA and drugs, exp(·) is the natural exponential function, and Y(l i )、Y(l j )、Y(d i ) and Y(d j ) are lncRNA l i 、lncRNA l j , drugs d i and drugs j The eigenvector representation of ||Y(l i )-Y(l j )|| 2 lncRNA l i and lncRNA l j The square of the Euclidean distance between the eigenvectors of ||Y(d i )-Y(d j )|| 2 Indicates drug d i With drug d j The squared Euclidean distance between the eigenvectors of .
[0147] The probability acquisition module includes:
[0148]
[0149] in, represents the output features of the lth layer, represents the output feature of the l-1 layer, W (l) is the weight matrix of the lth layer, b (l) is the bias term of the lth layer, and ReLU(·) is the activation function.
[0150] The probability acquisition module also includes: training the LDA-GNN model and LDA-DHG model embedded with the constraint strategy:
[0151] The lncRNA-drug association data, the lncRNA similarity matrix after dimensionality reduction, and the drug similarity matrix are input as edges and nodes of the graph structure into the LDA-GNN model built based on GAT and GCN. The node update method is as follows:
[0152]
[0153] in, Represents the feature vector of node j at layer 0 in the LDA-GNN model, represents the feature vector of node i in the first layer of the LDA-GNN model, d(i) represents the degree of node i, d(j) represents the degree of node j, N(i) represents the set of all neighbors of node i, and w 1 Represents the parameter matrix that GCN needs to learn;
[0154] The lncRNA-drug association data, the lncRNA similarity matrix after dimensionality reduction, and the drug similarity matrix are input as edges and nodes of the graph structure into the LDA-DHG model built based on GAT and GCN. The node update method is as follows:
[0155]
[0156] in, represents the feature of node i at the lth layer in the LDA-DHG model, Represents the feature of node i in the l+1th layer in the LDA-DHG model, Agg is the aggregation operation, and MLP is the multi-layer perceptron used for feature update;
[0157] Introduce the attention mechanism and calculate the attention coefficient between nodes:
[0158]
[0159] Among them, e ij is the attention coefficient of node i to neighbor node j, is the feature representation of node i in the kth layer, is the feature representation of node j in the kth layer, a is the shared attention mechanism for calculating the attention coefficient; W (k) represents the parameter matrix that needs to be learned in the kth layer, and σ represents the activation function;
[0160] Use the Softmax function to normalize the attention coefficient:
[0161]
[0162] Among them, a ij is the normalized value of the attention coefficient, e im For node i to all its neighbor nodes m∈Ni The attention coefficient;
[0163] Weighted sum:
[0164]
[0165] The probability acquisition module also includes:
[0166]
[0167] Among them, y i represents the true label of the i-th sample, y i ' represents the predicted value of the i-th sample, x a represents the embedding vector of the a-th associated pair, y b represents the embedding vector of the b-th associated pair, α and β are hyperparameters, N is the different classification results, and n is the dimension of each vector.
[0168] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. A drug-lncRNA relationship prediction method based on embedded constraints, characterized by: The method comprises: S1. Collect and preprocess lncRNA-drug association data to construct a dataset; S2. Based on the data set, construct a lncRNA-drug association matrix and extract a lncRNA-drug bipartite graph; S3. Based on the lncRNA-drug association matrix, obtaining a lncRNA similarity matrix and a drug similarity matrix, and performing dimensionality reduction processing on the lncRNA similarity matrix and the drug similarity matrix; S4. Inputting the lncRNA-drug bipartite graph, the lncRNA similarity matrix after dimensionality reduction, and the drug similarity matrix into an LDA-GNN model and an LDA-DHG model with an embedded constraint strategy to obtain a probability score of the association between the lncRNA and the drug; S5. Based on the probability score, predict the relationship between drug and lncRNA.
2. The method for predicting drug-lncRNA relationships based on embedded constraints according to claim 1, characterized in that: The S3 includes: G l (l i ,l j )=exp(-r l ||Y(l i )-Y(l j )|| 2 ) G d (d i ,d j )=exp(-r d ||Y(d i )-Y(d j )|| 2 ) Among them, G l (l i ,l j ) and G d (d i ,d j ) are lncRNA l i and lncRNA l j Gaussian similarity and drug d i With drug d j Gaussian similarity, r l and r d are the scale functions of the Gaussian kernel function in controlling the similarity between lncRNA and drugs, exp(·) is the natural exponential function, and Y(l i )、Y(l j )、Y(d i ) and Y(d j ) are lncRNA l i 、lncRNA l j , drugs d i and drugs j The eigenvector representation of ||Y(l i )-Y(l j )|| 2 lncRNAl i and lncRNAl j The square of the Euclidean distance between the eigenvectors of ||Y(d i )-Y(d j )|| 2 Indicates drug d i With drug d j The squared Euclidean distance between the eigenvectors of .
3. The method for predicting drug-lncRNA relationships based on embedded constraints according to claim 1, characterized in that: The S4 includes: in, represents the output features of the lth layer, represents the output feature of the l-1 layer, W (l) is the weight matrix of the lth layer, b (l) is the bias term of the lth layer, and ReLU(·) is the activation function.
4. The method for predicting drug-lncRNA relationships based on embedded constraints according to claim 3, characterized in that: The S4 further includes: training the LDA-GNN model and the LDA-DHG model with the embedding constraint strategy: The lncRNA-drug association data, the lncRNA similarity matrix after dimensionality reduction, and the drug similarity matrix are input as edges and nodes of the graph structure into the LDA-GNN model constructed based on GAT and GCN, and the node update method is as follows: in, Represents the feature vector of node j at layer 0 in the LDA-GNN model, represents the feature vector of node i in the first layer of the LDA-GNN model, d(i) represents the degree of node i, d(j) represents the degree of node j, N(i) represents the set of all neighbors of node i, and w 1 Represents the parameter matrix that GCN needs to learn; The lncRNA-drug association data, the lncRNA similarity matrix after dimensionality reduction, and the drug similarity matrix are input as edges and nodes of the graph structure into the LDA-DHG model constructed based on GAT and GCN, and the node update method is as follows: in, represents the feature of node i at the lth layer in the LDA-DHG model, Represents the feature of node i in the l+1th layer in the LDA-DHG model, Agg is the aggregation operation, and MLP is the multi-layer perceptron used for feature update; Introduce the attention mechanism and calculate the attention coefficient between nodes: Among them, e ij is the attention coefficient of node i to neighbor node j, is the feature representation of node i in the kth layer, is the feature representation of node j in the kth layer, a is the shared attention mechanism for calculating the attention coefficient; W (k) represents the parameter matrix that needs to be learned in the kth layer, and σ represents the activation function; Use the Softmax function to normalize the attention coefficient: Among them, a ij is the normalized value of the attention coefficient, e im For node i to all its neighbor nodes m∈N i The attention coefficient; Weighted sum:
5. The method for predicting drug-lncRNA relationships based on embedded constraints according to claim 4, characterized in that: Said S4 further comprises: Among them, y i represents the true label of the i-th sample, y i ' represents the predicted value of the i-th sample, x a represents the embedding vector of the a-th associated pair, y b represents the embedding vector of the b-th associated pair, α and β are hyperparameters, N is the different classification results, and n is the dimension of each vector.
6. A drug-lncRNA relationship prediction system based on embedded constraints, used to implement the drug-lncRNA relationship prediction method based on embedded constraints according to any one of claims 1 to 5, characterized in that: The system includes: an acquisition module, a bipartite graph extraction module, a similarity matrix acquisition module, a probability acquisition module and a prediction module; The acquisition module is used to collect lncRNA-drug association data and perform preprocessing to construct a data set; The bipartite graph extraction module is used to construct a lncRNA-drug association matrix based on the data set and extract a lncRNA-drug bipartite graph; The similarity matrix acquisition module is used to obtain a lncRNA similarity matrix and a drug similarity matrix based on the lncRNA-drug association matrix, and perform dimensionality reduction processing on the lncRNA similarity matrix and the drug similarity matrix; The probability acquisition module is used to input the lncRNA-drug bipartite graph and the lncRNA similarity matrix and the drug similarity matrix after dimensionality reduction into the LDA-GNN model and the LDA-DHG model with an embedded constraint strategy to obtain the probability score of the association between lncRNA and drug; The prediction module is used to predict the relationship between drug and lncRNA based on the probability score.
7. The drug-lncRNA relationship prediction system based on embedded constraints according to claim 6, characterized in that: The similarity matrix acquisition module includes: G l (l i ,l j )=exp(-r l ||Y(l i )-Y(l j )|| 2 ) G d (d i ,d j )=exp(-r d ||Y(d i )-Y(d j )|| 2 ) Among them, G l (l i ,l j ) and G d (d i ,d j ) are lncRNA l i and lncRNA l j Gaussian similarity and drug d i With drug d j Gaussian similarity, r l and r d are the scale functions of the Gaussian kernel function in controlling the similarity between lncRNA and drugs, exp(·) is the natural exponential function, and Y(l i )、Y(l j )、Y(d i ) and Y(d j ) are lncRNA l i 、lncRNA l j , drugs d i and drugs j The eigenvector representation of ||Y(l i )-Y(l j )|| 2 lncRNAl i and lncRNAl j The square of the Euclidean distance between the eigenvectors of ||Y(d i )-Y(d j )|| 2 Indicates drug d i With drug d j The squared Euclidean distance between the eigenvectors of .
8. The drug-lncRNA relationship prediction system based on embedded constraints according to claim 6, characterized in that: The probability acquisition module includes: in, represents the output features of the lth layer, represents the output feature of the l-1 layer, W (l) is the weight matrix of the lth layer, b (l) is the bias term of the lth layer, and ReLU(·) is the activation function.
9. The drug-lncRNA relationship prediction system based on embedded constraints according to claim 8, characterized in that: The probability acquisition module further includes: training the LDA-GNN model and the LDA-DHG model with the embedded constraint strategy: The lncRNA-drug association data, the lncRNA similarity matrix after dimensionality reduction, and the drug similarity matrix are input as edges and nodes of the graph structure into the LDA-GNN model constructed based on GAT and GCN, and the node update method is as follows: in, Represents the feature vector of node j at layer 0 in the LDA-GNN model, represents the feature vector of node i in the first layer of the LDA-GNN model, d(i) represents the degree of node i, d(j) represents the degree of node j, N(i) represents the set of all neighbors of node i, and w 1 Represents the parameter matrix that GCN needs to learn; The lncRNA-drug association data, the lncRNA similarity matrix after dimensionality reduction, and the drug similarity matrix are input as edges and nodes of the graph structure into the LDA-DHG model constructed based on GAT and GCN, and the node update method is as follows: in, represents the feature of node i at the lth layer in the LDA-DHG model, Represents the feature of node i in the l+1th layer in the LDA-DHG model, Agg is the aggregation operation, and MLP is the multi-layer perceptron used for feature update; Introduce the attention mechanism and calculate the attention coefficient between nodes: Among them, e ij is the attention coefficient of node i to neighbor node j, is the feature representation of node i in the kth layer, is the feature representation of node j in the kth layer, a is the shared attention mechanism for calculating the attention coefficient; W (k) represents the parameter matrix that needs to be learned in the kth layer, and σ represents the activation function; Use the Softmax function to normalize the attention coefficient: Among them, a ij is the normalized value of the attention coefficient, e im For node i to all its neighbor nodes m∈N i The attention coefficient; Weighted sum:
10. The drug-lncRNA relationship prediction system based on embedded constraints according to claim 9, characterized in that: The probability acquisition module also includes: Among them, y i represents the true label of the i-th sample, y i ' represents the predicted value of the i-th sample, x a represents the embedding vector of the a-th associated pair, y b represents the embedding vector of the b-th associated pair, α and β are hyperparameters, N is the different classification results, and n is the dimension of each vector.
Citation Information
Patent Citations
Drug-disease association prediction method and system
CN113140327A
IncRNA and disease association prediction method fusing heterogeneous network and graph neural network
CN114093425A
Drug relevance prediction method and device, terminal equipment and medium
CN116129989A
Characteristic fusion network-based ncRNA-drug resistance association prediction method
CN118248208A
Method for predicting ncRNA-disease-drug potential correlation degree
CN119008035A