Drug disease association prediction method based on curvature optimization association
By optimizing the drug-disease association graph and introducing similarity-perceptual hyperbolic graph neural network and gated signal self-attention mechanism, the problem of excessive compression and insufficient self-attention mechanism in drug-disease association prediction is solved, and the accuracy of prediction and generalization ability of the model are improved.
Patent Information
- Application Number
- CN202510430927.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-08
AI Technical Summary
In the existing drug-disease association prediction methods, the graph neural network has an overcompression problem, which limits the capture ability of long-range dependencies, the self-attention mechanism does not fully consider the semantics of the disease, and the neighborhood weight calculation ignores the impact of node similarity.
By designing the curvature optimization association modification module, the drug-disease association graph is optimized, and the similarity-aware hyperbolic graph neural network and the gating signal self-attention mechanism are used to improve feature representation accuracy and model performance.
It effectively alleviates the problem of overcompression of graph neural networks, improves the accuracy of feature representation of drug disease association prediction and the generalization ability of model.
Smart Images

Figure CN120388650A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of drug development, and specifically to a method for predicting drug-disease associations based on curvature optimization correlation. Background Art
[0002] Traditional drug discovery is a time-consuming and labor-intensive process. It usually takes more than a decade from drug discovery to clinical application, and the cost is high, usually between $500 million and $2 billion. With the rapid development of artificial intelligence technology, drug repositioning (DR) has received extensive attention as an alternative. Drug repositioning has significant advantages in accelerating drug discovery and reducing duplicate work, and can greatly shorten the drug development time and reduce capital waste. Drug repositioning, also known as drug repurposing, refers to the discovery of new uses for existing drugs, especially for the treatment of new diseases or symptoms. This practice is of great significance in the field of drug development, and its main method is to predict potential drug-disease associations.
[0003] Many researchers use computational methods to predict potential drug-disease associations in a more efficient and cost-effective manner. Significant progress has been made in DR research based on computational methods, which can identify new drug-disease associations (DDAs). These methods generally include the following two steps. First, relevant features are extracted from biochemical and medical data related to drugs and diseases. Then, these features are input into a trained model to facilitate DDA prediction. Such methods are mainly divided into two categories: machine learning (ML)-based methods and deep learning (DL)-based methods.
[0004] ML-based methods can extract features through various algorithms (such as matrix factorization, multi-kernel learning, support vector machines, and neural networks, etc.), and use these extracted embedded features to predict new DDAs. However, ML-based methods rely on shallow features and have limited representation ability for more advanced features. To overcome this limitation, DL-based methods (such as graph neural networks (GNNs), hyperbolic graph neural networks, and recurrent neural networks, etc.) have been developed to utilize their powerful representation learning to obtain the latent features of drugs and diseases.
[0005] However, in current research, the common way for DL-based methods to extract drug-disease related features is to construct a drug-disease association heterogeneous network from biomedical-related data sources, and then use GNNs or hyperbolic graph neural networks, etc. to extract features from the network. However, GNNs have the problem of over-compression, which affects the ability of message passing between graph nodes. Since GNNs can only pass messages at one distance per layer, this limits its ability to capture long-range dependencies.
[0006] In hyperbolic graph neural networks, calculating neighborhood weights is a crucial step as it determines how to aggregate information in the hyperbolic space. Common ways to calculate neighborhood weights include those based on attention mechanisms, curvature-based weights, and distance-based weights. Distance-based weights refer to calculating weights according to the hyperbolic distance between nodes in the hyperbolic space. The closer the nodes are, the greater the weights, which can be achieved through the reciprocal of the hyperbolic distance or its function. However, this approach ignores the impact of node similarity on the network.
[0007] In drug-disease association prediction, the self-attention mechanism is also often used to extract the features of drugs and diseases. The self-attention mechanism, also known as intra-attention, is a deep learning technique that can enhance the model's ability to process sequential data. It allows the model to dynamically focus on information at different positions in the sequence, thereby capturing the dependencies within the sequence. The semantic value of a disease usually refers to the richness of information or the level of detail related to a specific disease. In drug-disease association prediction, the semantic value of a disease is of great significance. Therefore, it is of great importance to improve the attention of the self-attention mechanism to the semantic value.
[0008] Based on the above, the current methods have the following drawbacks:
[0009] 1. For GNNs, over-compression is a common problem, which affects the graph's ability to transmit messages between nodes; since GNNs can only transmit messages at a distance of 1 per layer, this limits the GNN's ability to capture long-range dependencies;
[0010] 2. For hyperbolic graph neural networks, calculating neighborhood weights is a key step; however, existing calculation methods often ignore the impact of node similarity on the network;
[0011] 3. The traditional self-attention mechanism allows the model to dynamically focus on information at different positions in the sequence when processing sequential data, thereby capturing the dependencies within the sequence; however, for drug-disease association prediction, the semantics of diseases contribute more significantly to the association score; therefore, it is necessary to optimize the traditional self-attention mechanism to better capture disease semantic information.
[0012] Based on this, a drug-disease association prediction method based on curvature-optimized association is invented. Summary of the Invention
[0013] To solve the above technical problems, according to one aspect of the present invention, the present invention provides the following technical solutions:
[0014] A drug-disease association prediction method based on curvature-optimized association, which includes the following specific steps:
[0015] S1. Obtain the drug-disease association heterogeneous graph A, the drug similarity matrix DR, and the disease similarity matrix DS:
[0016] S11: Obtain drug-disease association data from existing public databases, construct a drug-disease association heterogeneous graph, and represent it as a matrix where m and d represent the numbers of drugs and diseases respectively;
[0017] S12: Obtain the chemical structures of drugs from existing public databases, and construct a drug similarity matrix
[0018] S13: Obtain disease phenotype information from existing public databases, measure disease similarity by disease phenotype similarity, and construct a disease similarity matrix
[0019] S2. Design a curvature-optimized association modification module to optimize the drug-disease association heterogeneous graph A to obtain the drug-disease association heterogeneous graph M * :
[0020] S21: Design a curvature-optimized association modification module to alleviate the over-compression problem in GCN;
[0021] S22: The drug-disease association heterogeneous graph A is processed by the association modification module to obtain the optimized drug-disease association heterogeneous graph M * ;
[0022] S3. Use GCN to perform feature extraction on the drug-disease association heterogeneous graph M * to obtain the first feature matrix F1;
[0023] S4. Design a similarity-aware hyperbolic graph neural network model to perform feature extraction on the matrix F1 to obtain the second feature matrix F2;
[0024] S5. Design a gated signal self-attention mechanism to perform feature extraction on the feature matrix F2 to obtain the third feature matrix F3;
[0025] S6. Based on the feature matrix F3, select MLP as the decoder and output the association prediction score S.
[0026] As a preferred solution of the drug-disease association prediction method based on curvature-optimized association described in the present invention, wherein: the definition of A in S11 is as follows:
[0027]
[0028] As a preferred solution of the drug-disease association prediction method based on curvature-optimized association described in the present invention, wherein: the specific steps of S12 are as follows:
[0029] S121: After obtaining the chemical structure of the drug from the existing public database, use the RDKit toolkit to convert the chemical structure of the drug into the topological fingerprint of the molecule;
[0030] S122: Calculate the Tanimoto similarity coefficient for each pair of drugs;
[0031] The calculation formula of the Tanimoto similarity coefficient is:
[0032]
[0033] where, R A and R B respectively represent the sets of two molecular fingerprints, and store the calculated Tanimoto similarity coefficient of each pair of drugs in the matrix DR. The finally obtained DR matrix is the drug similarity matrix.
[0034] As a preferred scheme of a drug-disease association prediction method based on curvature optimization association described in the present invention, wherein: the specific steps of S13 are as follows:
[0035] S131: Obtain disease phenotype information from the existing public database and structure the disease phenotype information using the Human Phenotype Ontology, which provides a hierarchical relationship between diseases and phenotypes;
[0036] S132: Based on the Human Phenotype Ontology, the phenotypic similarity between two diseases can be calculated. For this purpose, the information content method can be used to measure the specificity of phenotypic terms, and then the similarity between diseases can be calculated. The specific formula is as follows:
[0037]
[0038] where, d i and d j represent diseases, T(d i ) and T(d j ) respectively represent the sets of phenotypic terms of diseases d i and d j in the Human Phenotype Ontology, t represents a phenotypic term, IC(t) represents the information content of the phenotypic term t, store the phenotypic similarity values of all disease pairs in the matrix DS, and the element DS(i,j) in the matrix represents the similarity between diseases d i and d j .
[0039] As a preferred scheme of a drug-disease association prediction method based on curvature optimization association described in the present invention, wherein: the execution steps of the association modification module in S2 are as follows:
[0040] S211: Construct a new associated heterogeneous network M using DS, DR, and the drug-disease associated heterogeneous graph A, i.e.:
[0041]
[0042] S212: Calculate the Ollivier-Ricci curvature of all links in the entire drug-disease associated heterogeneous graph through M, and find the n a edges with the smallest Ollivier-Ricci curvature values, and add virtual similarity links to these edges to alleviate the negative curvature problem. At the same time, find the n r associated edges with the largest Ollivier-Ricci curvature values, and reduce the positive curvature in the network by removing them. Finally, retain the edges connecting the drug nodes and the disease nodes to obtain the optimized drug-disease associated graph M * :
[0043]
[0044] where A * is the matrix obtained by removing the large curvature edges and adding virtual similarity links to the drug-disease associated heterogeneous graph A, and A *T is the transpose matrix of A * .
[0045] As a preferred solution of the drug-disease association prediction method based on curvature optimization association described in the present invention, wherein: the specific steps of S3 are as follows:
[0046] Use GCN to perform feature extraction on the drug-disease associated graph M * to obtain the feature matrix F1, and define the formula:
[0047]
[0048] where, where is the identity matrix, is the degree matrix of the drug-disease associated graph, which is a diagonal matrix, represents taking the reciprocal of each diagonal element of the degree matrix and then taking the square root, represents the training parameter of the l-th layer of GCN, LeakyReLU represents the non-linear activation function, and the input layer H (1) is M * , and L is the number of convolutional layers.
[0049] As a preferred solution of the drug-disease association prediction method based on curvature optimization association described in the present invention, wherein: the specific steps of S4 are as follows:
[0050] S41: Use exponential mapping to map the Euclidean feature vector of each node in the feature matrix f1 of drug-disease associations to a hyperbolic space feature vector; for the drug-disease feature matrix F1, x represents a point in the hyperbolic space, represents the Euclidean feature vector of node x, and the mapping formula is as follows:
[0051]
[0052] where denotes the feature vector after being mapped to the hyperbolic space, represents the exponential mapping starting from x; c is the curvature of the hyperbolic space, and x represents a point in the hyperbolic space as the starting point of the exponential mapping; denotes the Möbius addition in the Poincaré ball model, which is a distance metric in the hyperbolic space; tanh represents the hyperbolic tangent function used to map real numbers to the interval (-1, 1); is the absolute value of the square root of the curvature c, used to adjust the scale of the mapping; is a scaling factor related to point x, which depends on the curvature c and the position of point x in the hyperbolic space; represents the feature vector 's Euclidean norm, representing the length of the feature vector; denotes the normalization operation, normalizing the feature vector and adjusting it according to the curvature c. Its formula maps the feature vector of drug-disease associations in the Euclidean space to the feature vector in the hyperbolic space;
[0053] S42: Calculate the similarity-aware neighborhood weight α ij , making the feature aggregation of the drug-disease features by the hyperbolic graph neural network robust. The calculation formula of α ij is as follows:
[0054]
[0055] where d i is the degree of node i;
[0056] S43: Perform feature aggregation on the features of drug-disease to obtain the final output feature representation After calculating the neighborhood weight, the features of drug-disease can be feature-aggregated. In the hyperbolic space, feature aggregation can be achieved in the following way:
[0057]
[0058] where L represents the number of layers, and l starts from 1, is the feature representation of node i at the l-th layer, is the neighbor set of node i, is the Möbius addition in the hyperbolic space, is the aggregation operation on neighbor features, α ij is the weight between nodes i and j, ⊙ represents element-wise multiplication, and the final output feature representation is
[0059] S44: Apply a non-linear activation function to the feature to introduce non-linearity; introduce the hyperbolic tangent function tanh as the non-linear activation function to introduce non-linearity, thereby enhancing the expressive power of the model. The defined formula is:
[0060]
[0061] where, is the representation after passing through the non-linear function;
[0062] S45: Use logarithmic mapping to map the feature vector of drug-disease in the hyperbolic space back to the feature vector in the Euclidean space:
[0063]
[0064] where, represents the logarithmic mapping starting from q; q is a point in the hyperbolic space, and generally the origin, i.e., the zero vector, is selected as the reference point for the logarithmic mapping; is another point in the hyperbolic space, and the goal is to find the tangent vector from q to ; c is the curvature of the hyperbolic space; is a scaling factor depending on the curvature c and the position of point q, specifically represents the Möbius addition in the Poincaré ball model; tanh -1 is the inverse function of the hyperbolic tangent function, i.e., the inverse hyperbolic tangent function; ||·||2 represents the Euclidean norm. The geometric meaning of its formula is that given two points q and in the hyperbolic space, the logarithmic mapping calculates the tangent vector of the geodesic from q to at q, and the tangent vector can be regarded as the Euclidean representation of the "direction" and "distance" between q and . The feature matrix mapped back to the Euclidean space is represented as
[0065] As a preferred scheme of a drug-disease association prediction method based on curvature optimization correlation described in the present invention, wherein: The specific steps of S5 are as follows:
[0066] S51: A gating mechanism is introduced, enabling the self-attention mechanism to pay more attention to the drug-disease features that are more important for the current drug redirection task when processing drug-disease information, thereby improving the performance and generalization ability of the model. The defining formula is:
[0067] Q = F2W Q #(9)
[0068] K = F2W K #(10)
[0069] V = F2W V #(11)
[0070] Among them, and are learnable weight matrices;
[0071] S52: It is necessary to calculate a gating signal, which will be used to adjust the contribution of the value matrix V. The gating signal can be calculated through the dot product of the query matrix Q and the key matrix K, and then generated through a learnable weight matrix and a non-linear activation function sigmoid;
[0072] GateSignal = sigmoid(QK T W G )#(12)
[0073] S53: Applying the gating signal to the value matrix V can be achieved by multiplying the gating signal by the value matrix V, thereby controlling the weight of each value vector in the final output;
[0074] GatedValues = GateSignal × V#(13)
[0075]
[0076] where d k is the dimension of the vector K, the Softmax function performs the normalization operation, σ represents the ReLU operation, and GateSignal represents the gating signal.
[0077] As a preferred solution of the drug-disease association prediction method based on curvature optimization association described in the present invention, wherein: the specific steps of S6 are as follows:
[0078] Based on the feature matrix F3, where Z i,j represents the drug-disease association feature of drug i and disease j in the feature matrix F3, S i,j is drug m i and disease d jFor the association score between them, a multi-layer linear perceptron MLP is selected as the decoder to predict the probability of each category:
[0079]
[0080] where and represent learnable parameters, b r and b s are bias terms, and σ is an optional activation function.
[0081] Compared with the prior art:
[0082] The association modification module based on curvature optimization in the present invention can optimize the drug-disease association graph. The optimized drug-disease association graph can effectively alleviate the over-compression problem in the GCN process and improve the accuracy of feature representation. When the similarity-aware hyperbolic graph neural network model extracts drug-disease features, it not only considers whether nodes are connected but also the similarity degree of the features of the nodes themselves, making the calculated neighborhood weights more robust. The gated signal self-attention mechanism calculates a gated signal, which will be used to adjust the contribution of the value matrix V, so that the gated signal self-attention mechanism can pay more attention to the features that are more important for the drug redirection task when processing information, thereby improving the performance and generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] Figure 1 is a schematic flow diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0084] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0085] The present invention provides a method for predicting drug-disease associations based on curvature-optimized associations. Please refer to Figure 1 for the specific steps as follows:
[0086] S1. Obtain a drug-disease association heterogeneous graph A, a drug similarity matrix DR, and a disease similarity matrix DS:
[0087] S11: Obtain drug-disease association data from existing public databases, construct a drug-disease association heterogeneous graph and represent it as a matrix where m and d represent the numbers of drugs and diseases respectively; among them, the definition of A in S11 is as follows:
[0088]
[0089] S12: Obtain the chemical structures of drugs from existing public databases and construct a drug similarity matrix
[0090] Among them, the specific steps of S12 are as follows:
[0091] S121: After obtaining the chemical structure of the drug from the existing public database, use the RDKit toolkit to convert the chemical structure of the drug into the topological fingerprint of the molecule;
[0092] S122: Calculate the Tanimoto similarity coefficient for each pair of drugs;
[0093] The calculation formula of the Tanimoto similarity coefficient is:
[0094]
[0095] Among them, R A and R B respectively represent the sets of two molecular fingerprints. Store the calculated Tanimoto similarity coefficient for each pair of drugs in the matrix DR. The finally obtained DR matrix is the drug similarity matrix
[0096] S13: Obtain the disease phenotype information from the existing public database, measure the disease similarity by the disease phenotype similarity, and construct the disease similarity matrix
[0097] Among them, the specific steps of S13 are as follows:
[0098] S131: Obtain the disease phenotype information from the existing public database and structure the disease phenotype information using the Human Phenotype Ontology, which provides the hierarchical relationship between diseases and phenotypes;
[0099] S132: Based on the Human Phenotype Ontology, the phenotype similarity between two diseases can be calculated. For this purpose, the information content method can be used to measure the specificity of phenotype terms, and then the similarity between diseases can be calculated. The specific formula is as follows:
[0100]
[0101] Among them, d i and d j represent diseases, T(d i ) and T(d j ) respectively represent the sets of phenotype terms of diseases d i and d j in the Human Phenotype Ontology. t represents the phenotype term, OC(t) represents the information content of the phenotype term t. Store the phenotype similarity values of all disease pairs in the matrix DS. The element DS(i, j) in the matrix represents the similarity between diseases d i and d j ;
[0102] S1 includes but is not limited to the following embodiments:
[0103] 18,416 verified drug-disease association relationships obtained from the DrugBank database, representing the drug-disease association relationships as a drug-disease association heterogeneous graph A; measuring the similarity of drugs by the similarity of drug molecular fingerprints, obtaining the drug fingerprint similarity from the DrugBank database based on RDKit and Tanimoto, and representing the drug fingerprint similarity matrix as DR; measuring the similarity of diseases by disease phenotype similarity, representing the disease similarity matrix as DS, and calculating the disease phenotype similarity by text mining and analyzing the biomedical information in the OMIM database;
[0104] Among them, (1) RDKit can generate different types of fingerprints for molecules, which is crucial for comparing the similarity between molecules; fingerprints are a way to encode the chemical structure information of molecules into binary or vectors, which can be used for subsequent similarity calculation and search; the Tanimoto coefficient is a measure for measuring the similarity between two sets, and in chemoinformatics, it is often used to calculate the similarity between two molecular fingerprints;
[0105] (2) Both the DrugBank database and the OMIM database are publicly available databases;
[0106] The DrugBank database is a comprehensive bioinformatics and chemoinformatics database that combines detailed drug data (such as chemical information, pharmacological information, drug targets, etc.) and comprehensive drug target information, and is an important tool for drug research and development, clinical trials, and safety monitoring;
[0107] OMIM (Online Mendelian Inheritance in Man) is an online database focusing on the relationship between human genes and genetic diseases, maintained by the National Library of Medicine (NLM) of the United States; it provides information such as detailed descriptions of genetic diseases, gene mutations, clinical phenotypes, inheritance patterns, and related literature;
[0108] S2, designing a curvature-optimized association modification module to optimize the drug-disease association heterogeneous graph A to obtain the drug-disease association heterogeneous graph M * :
[0109] S21: Designing a curvature-optimized association modification module to alleviate the over-compression problem in GCN;
[0110] S22: The drug-disease association heterogeneous graph A is processed by the association modification module to obtain the optimized drug-disease association heterogeneous graph M * ;
[0111] Among them, the execution steps of the associated modification module in S2 are as follows:
[0112] S211: Use DS, DR, and the drug-disease associated heterogeneous graph A to construct a new associated heterogeneous network M, that is:
[0113]
[0114] S212: Calculate the Ollivier-Ricci curvature of all links in the entire drug-disease associated heterogeneous graph through M, and find the n a edges with the smallest Ollivier-Ricci curvature value, and add virtual similarity links to these edges to alleviate the negative curvature problem. At the same time, find the n r associated edges with the largest Ollivier-Ricci curvature value, and reduce the positive curvature in the network by removing them. Finally, retain the edges connecting the drug nodes and the disease nodes to obtain the optimized drug-disease associated graph M * :
[0115]
[0116] Among them, A * is the matrix obtained by removing the large curvature edges and adding virtual similarity links to the drug-disease associated heterogeneous graph A. A *T is the transpose matrix of A * ;
[0117] Among them, the over-compression problem refers to the occurrence of multi-layer propagation in the graph neural network, resulting in node features becoming too similar during the information mixing process, thus losing the important differences between nodes and the complex structural characteristics of the graph. This problem will reduce the model's expression ability and prediction accuracy for graph data; and since only 1 distance of messages can be transmitted per layer, this means that two nodes with a distance of K can only transmit information to each other when the GNN has at least K layers; this phenomenon limits the ability of the GNN to capture long-range dependencies;
[0118] The phenomenon of over-compression originates from the local structural features of the graph; the local structural features of the graph refer to the structural properties of a subgraph formed by a node or a small part of nodes and their adjacent nodes in the graph. These properties can describe the topological structure of the graph and the relationships between nodes within a local range, corresponding to the global structural features of the entire graph; it mainly focuses on the patterns and properties exhibited by the direct neighborhood of nodes in the graph and the set of nodes closely connected to it; at the same time, the positive and negative curvatures in the graph also affect the information transmission process on the graph to a certain extent; therefore, the characteristics of the local graph can be used to optimize the performance of the graph neural network (GNN); optimizing the GNN using graph curvature can improve the scale and performance of the graph structure to a certain extent; in the case where traditional curvature theory is applied to discrete graphs, its concept is no longer valid, so the concept of Ollivier-Ricci curvature is introduced;
[0119] The Ollivier-Ricci curvature takes into account the cost that a node has to pay when migrating to other nodes according to a specific probability distribution; define the transition probability μ u (x) as follows:
[0120]
[0121] where V is the nodes of the heterogeneous graph, x represents the neighbor nodes of node u, denoted as x ∈ N u , deg(u) refers to the degree of node u, and u ∼ x represents that there is a direct connection between two nodes; then the Wasserstein distance between two nodes u and v can be calculated as:
[0122]
[0123] where δ(x,y) represents the joint probability density between two nodes, and d(x,y) measures the geodesic distance between two points on the manifold; the Wasserstein distance W(μ u , μ v ) measures the cost that node u has to move when reaching node v through random walk;
[0124] According to the definition of Ollivier-Ricci, the Ollivier-Ricci curvature between nodes u and v is finally calculated as:
[0125]
[0126] Introduce the concept of virtual similarity links to find the links that truly affect message passing blockage; for two nodes that are not directly connected but have a high similarity, it can be considered that there is a virtual link of similarity link between them, which has an equivalent effect to the directly connected link in message passing; therefore, for each disease in DS, select m most similar diseases to construct virtual disease similarity links, and for each drug in DR, select m most similar drugs to construct virtual drug similarity links; find the links that still maintain negative curvature after adding virtual similarity links, and these links can be considered as the reasons for message passing blockage; reducing the number of these links can alleviate the over-smoothing problem in the message passing process;
[0127] S3. Use GCN to perform feature extraction on the drug-disease association heterogeneous graph M * to obtain the first feature matrix F1;
[0128] Among them, the specific steps of S3 are as follows:
[0129] Use GCN to perform feature extraction on the drug-disease association graph M * to obtain the feature matrix F1, and define the formula:
[0130]
[0131] Among them, Among them is the identity matrix, is the degree matrix of the drug-disease association graph, which is a diagonal matrix, represents taking the reciprocal of each diagonal element of the degree matrix and then taking the square root, represents the training parameter of the l-th layer of GCN, LeakyReLU represents the non-linear activation function, and the input layer H (1) is M * , and L is the number of convolutional layers;
[0132] Among them, the core idea of GCN (Graph Neural Network) is to aggregate the neighbor information of nodes into the representation of the nodes themselves through convolutional operations, similar to the convolutional operations of traditional Convolutional Neural Networks (CNNs) on images;
[0133] S4. Design a similarity-aware hyperbolic graph neural network model to perform feature extraction on the matrix F1 to obtain the second feature matrix F2;
[0134] Among them, the specific steps of S4 are as follows:
[0135] S41: Use the exponential mapping to map the Euclidean feature vectors of each node in the drug-disease association feature matrix F1 into hyperbolic space feature vectors; for the drug-disease feature matrix F1, x represents a point in hyperbolic space, The Euclidean eigenvector representing node x has the following mapping formula:
[0136]
[0137] where denotes the eigenvector after being mapped to the hyperbolic space, represents the exponential map starting from x; c is the curvature of the hyperbolic space, and x represents a point in the hyperbolic space, serving as the starting point of the exponential map; denotes the Möbius addition in the Poincaré ball model, which is a distance metric in the hyperbolic space; tanh represents the hyperbolic tangent function, used to map real numbers to the interval (-1, 1); is the absolute value of the square root of the curvature c, used to adjust the scale of the mapping; is a scaling factor related to point x, which depends on the curvature c and the position of point x in the hyperbolic space; represents the eigenvector of the Euclidean norm, representing the length of the eigenvector; denotes the normalization operation, which normalizes the eigenvector and adjusts it according to the curvature c. Its formula maps the eigenvector of the drug-disease association in the Euclidean space to the eigenvector in the hyperbolic space;
[0138] S42: Calculate the similarity-aware neighborhood weight α ij , making the hyperbolic graph neural network robust to the feature aggregation of drug-disease features. The calculation formula of α ij is as follows:
[0139]
[0140] where d i is the degree of node i;
[0141] Among them, when traditional graph neural networks calculate neighborhood weights, they may only consider the structural information between nodes, such as the connection relationship or distance information between nodes. The similarity-aware neighborhood weight introduces the cosine similarity function, which can more comprehensively consider the intrinsic similarity between nodes, meaning that it not only considers whether nodes are connected but also the degree of feature similarity of the nodes themselves;
[0142] S43: Perform feature aggregation on the features of drug-disease to obtain the final output feature representation After calculating the neighborhood weights, the features of drug-disease can be aggregated. In the hyperbolic space, feature aggregation can be achieved in the following way:
[0143]
[0144] where \(L\) represents the number of layers, \(l\) starts from 1, is the feature representation of node \(i\) at the \(l\)-th layer, is the neighbor set of node \(i\), is the Möbius addition in the hyperbolic space, is the aggregation operation on neighbor features, \(\alpha\) ij is the weight between nodes \(i\) and \(j\), \(\odot\) represents element-wise multiplication, and the final output feature representation is
[0145] S44: Apply a non-linear activation function to the feature to introduce non-linearity; introduce the hyperbolic tangent function \(\tanh\) as the non-linear activation function to introduce non-linearity, thereby enhancing the expressive power of the model. The defined formula is:
[0146]
[0147] where, is the representation after passing through the non-linear function;
[0148] S45: Use the logarithmic map to map the feature vector of the drug-disease in the hyperbolic space back to the feature vector in the Euclidean space:
[0149]
[0150] where, represents the logarithmic map starting from \(q\); \(q\) is a point in the hyperbolic space, and generally the origin, i.e., the zero vector, is selected as the reference point for the logarithmic map; is another point in the hyperbolic space, and the goal is to find the tangent vector from \(q\) to ; \(c\) is the curvature of the hyperbolic space; is a scale factor depending on the curvature \(c\) and the position of point \(q\), specifically represents the Möbius addition in the Poincaré ball model; \(\tanh\) -1 is the inverse function of the hyperbolic tangent function, i.e., the inverse hyperbolic tangent function; \(\|\cdot\|_2\) represents the Euclidean norm. The geometric meaning of the formula is that given two points \(q\) and in the hyperbolic space, the logarithmic map calculates the tangent vector of the geodesic from \(q\) to at \(q\), and the tangent vector can be regarded as the Euclidean representation of the "direction" and "distance" between \(q\) and . The feature matrix mapped back to the Euclidean space is represented as
[0151] Among them, the hyperbolic graph neural network is a deep learning architecture that combines graph neural networks with hyperbolic geometry, aiming to utilize the characteristics of hyperbolic space to process graph data. Especially for data with hierarchical or tree-like structures, hyperbolic space can better represent its inherent geometric characteristics because hyperbolic space has higher expressive power than Euclidean space and can more effectively represent data with complex structures in lower dimensions;
[0152] S5. Design a gated signal self-attention mechanism to extract features from the feature matrix F2 to obtain the third feature matrix F3;
[0153] Among them, the specific steps of S5 are as follows:
[0154] S51: Introduce a gating mechanism so that the self-attention mechanism can pay more attention to those drug-disease features that are more important for the current drug redirection task when processing drug-disease information, thereby improving the performance and generalization ability of the model. Define the formula:
[0155] Q = F2W Q #(9)
[0156] K = F2W K #(10)
[0157] V = F2W V #(11)
[0158] Among them, and are learnable weight matrices;
[0159] S52: It is necessary to calculate a gating signal, which will be used to adjust the contribution of the value matrix V. The gating signal can be calculated through the dot product of the query matrix Q and the key matrix K, and then generated through a learnable weight matrix and a non-linear activation function sigmoid;
[0160] GateSignal = sigmoid(QK T W G ) #(12)
[0161] S53: Apply the gating signal to the value matrix V, which can be achieved by multiplying the gating signal by the value matrix V, thereby controlling the weight of each value vector in the final output;
[0162] GatedValues = GateSignal × V #(13)
[0163]
[0164] where d kis the dimension of vector K, the Softmax function performs a normalization operation, σ represents performing a ReLU operation, and GateSignal represents the gating signal;
[0165] Among them, the gating signal self-attention mechanism can dynamically weight different features according to the requirements of the drug redirection task by introducing a gating signal. This enables the model to pay more attention to the features that are more important for the current task when processing information;
[0166] S6. Based on the feature matrix F3, select MLP as the decoder to output the association prediction score S;
[0167] Among them, based on the feature matrix F3, where Z i,j represents the association feature between drug i and disease j in the feature matrix F3, and S i,j is the drug m i and the disease d j The association score between them, select the multi-layer linear perceptron MLP as the decoder to predict the probability of each category:
[0168]
[0169] where and represent learnable parameters, b r and b s are bias terms, and σ is an optional activation function.
[0170] In summary, the present invention designs an association modification module based on curvature optimization to alleviate the over-compression problem in GCN; specifically, calculate the Ollivier-Ricci curvature of all links in the entire drug-disease association heterogeneous graph, find the n a edges with the smallest Ollivier-Ricci curvature value, and add virtual similarity links to these edges to alleviate the negative curvature problem. At the same time, find the n r association edges with the largest Ollivier-Ricci curvature value, and reduce the positive curvature in the network by removing them; finally, retain the edges connecting the drug nodes and the disease nodes to obtain an optimized drug-disease association graph. The optimized drug-disease association graph can alleviate the over-compression problem in GCN;
[0171] The present invention designs a similarity-aware hyperbolic graph neural network model. By introducing a cosine similarity function when calculating the neighborhood weights, the model not only considers whether the nodes are connected when extracting drug-disease features, but also considers the degree of similarity of the features of the nodes themselves;
[0172] The present invention designs a gated signal self-attention mechanism. By calculating a gated signal, this gated signal will be used to adjust the contribution of the value matrix V, enabling the gated signal self-attention mechanism to pay more attention to the more important features of the drug redirection task when processing information, thereby improving the performance and generalization ability of the model.
[0173] Although the present invention has been described above with reference to the embodiments, various improvements can be made to it and components therein can be replaced with equivalents without departing from the scope of the present invention. In particular, as long as there is no structural conflict, the various features in the embodiments disclosed in the present invention can be combined with each other in any way, and the exhaustive description of these combinations is not given in this specification only for the sake of saving space and resources. Therefore, the present invention is not limited to the specific embodiments disclosed in the text, but includes all technical solutions falling within the scope of the claims.
Claims
1. A drug-disease association prediction method based on curvature optimization correlation, characterized in that The specific steps are as follows: S1. Obtain the drug-disease association heterogeneous graph A, the drug similarity matrix DR, and the disease similarity matrix DS: S11: Obtain drug-disease association data from existing public databases, construct a drug-disease association heterogeneous graph and represent it as a matrix where m and d represent the numbers of drugs and diseases respectively; S12: Obtain the chemical structure of the drug from the existing public database and construct a drug similarity matrix S13: Obtain disease phenotype information from existing public databases, measure disease similarity by disease phenotype similarity, and construct a disease similarity matrix S2. Design an associated modification module for curvature optimization to optimize the drug-disease association heterogeneous graph A and obtain the drug-disease association heterogeneous graph M * : S21: Design an association modification module with curvature optimization to alleviate the over-compression problem in GCN; S22: The drug-disease association heterogeneous graph A is processed by the association modification module to obtain the optimized drug-disease association heterogeneous graph M * ; S3. Use GCN to perform feature extraction on the drug-disease association heterogeneous graph M * to obtain the first feature matrix F1; S4. Design a similarity-aware hyperbolic graph neural network model to extract features from the matrix F1 to obtain the second feature matrix F2; S5. Design a gated signal self-attention mechanism to extract features from the feature matrix F2 to obtain the third feature matrix F3; S6. Based on the feature matrix F3, select MLP as the decoder and output the association prediction score S.
2. The drug-disease association prediction method based on curvature optimization association according to claim 1, wherein The definition of A in S11 is as follows:
3. A method for predicting drug-disease association based on curvature optimization correlation according to claim 1, characterized in that, The specific steps of S12 are as follows: S121: After obtaining the chemical structure of the drug from the existing public database, use the RDKit toolkit to convert the chemical structure of the drug into the topological fingerprint of the molecule; S122: Calculate the Tanimoto similarity coefficient for each pair of drugs; The calculation formula of the Tanimoto similarity coefficient is: where R A and R B respectively represent two sets of molecular fingerprints. The calculated Tanimoto similarity coefficients for each pair of drugs are stored in the matrix DR, and the finally obtained DR matrix is the drug similarity matrix.
4. The drug-disease association prediction method based on curvature optimization association according to claim 1, wherein The specific steps of S13 are as follows: S131: Obtain the disease phenotype information from the existing public database and structure the disease phenotype information using the Human Phenotype Ontology, which provides the hierarchical relationship between diseases and phenotypes; S132: Based on the Human Phenotype Ontology, the phenotypic similarity between two diseases can be calculated. For this purpose, the information content method can be used to measure the specificity of phenotypic terms, and then the similarity between diseases can be calculated. The specific formula is as follows: Among them, d i and d j represent diseases, T(d i ) and T(d j ) respectively represent the set of phenotypic terms of disease d i and d j in the Human Phenotype Ontology, t represents a phenotypic term, IC(t) represents the information content of the phenotypic term t, and the phenotypic similarity values of all disease pairs are stored in the matrix DS. The element DS(i, j) in the matrix represents the similarity between disease d i and d j .
5. A method for predicting the association between drugs and diseases based on curvature optimization association according to claim 1, wherein The execution steps of the association modification module in S2 are as follows: S211: Use DS, DR, and the drug-disease association heterogeneous graph A to construct a new association heterogeneous network M, that is: S212: Calculate the Ollivier-Ricci curvature of all links in the entire drug-disease association heterogeneous graph through M, find the n a edges with the smallest Ollivier-Ricci curvature values, and add virtual similarity links to these edges to alleviate the negative curvature problem. At the same time, find the n r associated edges with the largest Ollivier-Ricci curvature values, and reduce the positive curvature in the network by removing them. Finally, retain the edges connecting drug nodes and disease nodes to obtain the optimized drug-disease association graph M * : Where A * is the matrix obtained by removing the large-curvature edges and adding virtual similarity links to the drug-disease association heterogeneous graph A, and A *T is A * 's transpose matrix.
6. The method for predicting drug-disease association based on curvature optimization correlation according to claim 1, wherein The specific steps of S3 are as follows: Using GCN to perform feature extraction on the drug-disease association graph M * to obtain the feature matrix F1, and define the formula: Among them, Among them is the identity matrix, is the degree matrix of the drug-disease association graph, which is a diagonal matrix, represents taking the reciprocal of each diagonal element of the degree matrix and then taking the square root, represents the training parameter of the l-th layer GCN, LeakyReLU represents the non-linear activation function, and the input layer H( 1) is M * , and L is the number of layers of the convolutional layer.
7. A method for predicting the association between drugs and diseases based on curvature optimization association according to claim 1, wherein The specific steps of S4 are as follows: S41: Using exponential mapping, map the Euclidean feature vector of each node in the drug-disease association feature matrix F1 to a hyperbolic space feature vector; for the drug-disease feature matrix F1, x represents a point in the hyperbolic space, represents the Euclidean feature vector of node x, and the mapping formula is as follows: Among them, denotes the eigenvector after being mapped to the hyperbolic space, denotes the exponential map starting from x; c is the curvature of the hyperbolic space, and x represents a point in the hyperbolic space, serving as the starting point of the exponential map; denotes the Möbius addition in the Poincaré ball model, which is a distance metric in the hyperbolic space; tanh represents the hyperbolic tangent function, used to map real numbers to the interval (-1, 1); is the absolute value of the square root of the curvature c, used to adjust the scale of the mapping; is a scaling factor related to the point x, which depends on the curvature c and the position of the point x in the hyperbolic space; represents the eigenvector of the Euclidean norm, representing the length of the eigenvector; denotes the normalization operation, which normalizes the eigenvector and adjusts it according to the curvature c. Its formula maps the eigenvector of the drug-disease association in the Euclidean space to the eigenvector in the hyperbolic space; S42: Calculate the similarity-aware neighborhood weight α ij , making the hyperbolic graph neural network robust to the feature aggregation of drug-disease features. The calculation formula of α ij is as follows: where d i is the degree of node i; S43: Aggregate the features of the drug-disease to obtain the final output feature representation After calculating the neighborhood weights, it is possible to aggregate the features of the drug-disease. In the hyperbolic space, feature aggregation can be achieved in the following way: where L represents the number of layers, l starts from 1, is the feature representation of node i at the l-th layer, is the neighbor set of node i, is the Möbius addition in the hyperbolic space, is the aggregation operation on neighbor features, α ij is the weight between nodes i and j, ⊙ represents element-wise multiplication, and the final output feature representation is S44: Using a non-linear activation function as a feature Introduce non-linear characteristics; introduce the hyperbolic tangent function tanh as a non-linear activation function to introduce non-linear characteristics, thereby enhancing the expressive power of the model. The definition formula is: Among them, is the representation after passing through a non-linear function; S45: Use logarithmic mapping to map the feature vectors of drug diseases in the hyperbolic space back to the feature vectors in the Euclidean space: Among them, represents the logarithmic mapping starting from q; q is a point in the hyperbolic space, and it selects the origin, that is, the zero vector, as the reference point for the logarithmic mapping; is another point in the hyperbolic space, and the goal is to find the tangent vector from q to ; c is the curvature of the hyperbolic space; is a scale factor that depends on the curvature c and the position of the point q, specifically represents the Möbius addition in the Poincaré ball model; tanh -1 is the inverse function of the hyperbolic tangent function, that is, the inverse hyperbolic tangent function; ||·||2 represents the Euclidean norm, and the geometric meaning of its formula is that given two points q and in the hyperbolic space, the logarithmic mapping calculates the tangent vector of the geodesic from q to at q, and its tangent vector can be regarded as the Euclidean representation of the "direction" and "distance" between q and , and the characteristic matrix mapped back to the Euclidean space is represented as 8. A method for predicting drug-disease association based on curvature optimization association according to claim 1, characterized in that The specific steps of S5 are as follows: S51: Introduce a gating mechanism so that the self-attention mechanism can pay more attention to the drug-disease features that are more important for the current drug redirection task when processing drug-disease information, thereby improving the performance and generalization ability of the model. The definition formula is: Q = F2W Q #(9) K = F2W K #(10) V = F2W V #(11) Among them, and are learnable weight matrices; S52: It is necessary to calculate a gating signal, which will be used to adjust the contribution of the value matrix V. The gating signal can be calculated by querying the dot product of the query matrix Q and the key matrix K, and then generated through a learnable weight matrix and a non-linear activation function sigmoid; GateSignal=sigmoid(QK T W G )#(12) S53: Apply the gated signal to the value matrix V, which can be achieved by multiplying the gated signal by the value matrix V, thereby controlling the weight of each value vector in the final output; GatedValues = GateSignal × V#(13) where d k is the dimension of vector K, the Softmax function performs a normalization operation, σ represents the ReLU operation, and GateSignal represents the gating signal.
9. The method for predicting the association between drugs and diseases based on curvature optimization correlation according to claim 1, wherein The specific steps of S6 are as follows: Based on the feature matrix F3, where Z i,j represents the association feature between drug i and disease j in the feature matrix F3, and S i,j is the association score between drug m i and disease d j Select the multi-layer linear perceptron MLP as the decoder to predict the probability of each category: where and represent learnable parameters, b r and b s are bias terms, and σ is an optional activation function.
Citation Information
Patent Citations
System and method for predicting protein-ligand binding affinity based on filtering curvature
CN116312864A
Electronic medical record preprocessing method based on hyperbolic graph neural network
CN118136192A
Model training method, drug and disease association prediction method and related device
CN119541899A
Generating drug repositioning hypotheses based on integrating multiple aspects of drug similarity and disease similarity
US20160140312A1
Method and System for Assessing Drug Efficacy Using Multiple Graph Kernel Fusion
US20210134418A1