Drug and disease association prediction method based on hyperbolic graph feature learning network
The hyperbolic graph feature learning network extracts similar and heterogeneous features of drugs and diseases in hyperbolic space, solving the problem of simplified features and improper negative sample selection in the existing models, and achieving more accurate drug and disease association prediction.
Patent Information
- Application Number
- CN202510578790.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-05-07
AI Technical Summary
When existing drug and disease association prediction models learn the characteristics of drugs and diseases in the Euclidean space, there is an oversimplified similarity between nodes, which cannot accurately capture complex correlation patterns, limited feature representation capabilities, and improper negative sample selection affects model performance.
Using a hyperbolic graph feature learning network, similar and heterogeneous features are extracted by mapping drug and disease data to hyperbolic space, similar and heterogeneous features are extracted using hyperbolic graph feature restructuring and heterogeneous change graph converter, and high-quality feature matrix is generated through positive and negative fusion difficult-to-example sampling strategy for prediction.
Improve the accuracy and generalization ability of drug and disease association prediction, avoid distortion problems and negative sample noise in Euclidean space, and generate a comprehensive and high-quality feature representation.
Smart Images

Figure CN120565103A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of drug-disease association prediction, and in particular to a drug-disease association prediction method based on a hyperbolic graph feature learning network. Background Art
[0002] Drug repositioning (DR), also known as drug repurposing, is a promising approach in pharmaceutical innovation. It has attracted widespread attention from the academic community because it rapidly and cost-effectively discovers new clinical applications for approved drugs. The traditional drug discovery process typically takes 10 to 17 years and is costly, ranging from $500 million to $2 billion. It also takes at least a decade to bring a drug to market, a process that faces numerous challenges. In particular, with the rapid development of artificial intelligence, drug repositioning is becoming increasingly mainstream as an effective strategy to accelerate drug discovery. Repurposing approved drugs to discover new indications not only saves significant time and costs but also reduces the high risk of failure. Identifying new indications for existing drugs is crucial at all stages of drug discovery. In recent years, computational approaches driven by biomedical data have made significant progress in drug repositioning. These approaches accelerate the drug discovery process by extracting features from biochemical and medical data related to drugs and diseases and using machine learning (ML) and deep learning (DL) models to predict drug-disease associations (DDAs). However, many methods integrate information by simply splicing data from different modalities, which limits their ability to capture features from a comprehensive and in-depth perspective. Therefore, innovative computational methods are needed to better integrate multimodal information to improve the efficiency and accuracy of drug repositioning and further promote the advancement of drug discovery.
[0003] Graph Neural Network (GNN), as a deep learning (DL) method, is good at processing graph-structured data and can effectively capture the complex relationships and topological structures between nodes in drug-disease networks. Therefore, it has great potential in predicting the potential association between drugs and diseases. However, most of the existing GNN-based drug-disease association prediction models learn the characteristics of drugs and diseases in Euclidean space. For drug-drug homogeneous graphs and disease-disease homogeneous graphs, linear transformations and distance metrics in Euclidean space will oversimplify the similarity relationships between nodes, especially when the data presents a power-law distribution or significant hierarchical features. This simplification will cause the model to be unable to accurately capture the complex association patterns between nodes, ultimately affecting the accuracy of drug-disease association prediction. Drug-disease heterogeneous graphs usually have a tree-like structure, and their volume grows exponentially with the increase of radius, while the volume of a sphere in Euclidean space grows polynomially, which can easily lead to distortion.
[0004] Furthermore, negative samples play a crucial role in predicting drug-disease associations, but existing methods typically train by randomly sampling negative samples. This approach can overlook important negative samples or introduce irrelevant noise samples, which can affect model performance. Therefore, a method for drug-disease association prediction based on a hyperbolic graph feature learning network is proposed. This method not only more comprehensively understands the complex relationships between drug and disease nodes in the graph in hyperbolic space, but also synthesizes challenging negative samples, thereby further improving the accuracy of predictions of potential drug-disease associations. This method is not only innovative in theory but also has significant practical value.
[0005] Based on the above, the existing drug-disease association prediction model based on graph convolutional networks has the following problems:
[0006] 1. Existing drug-disease association prediction models typically learn similarity features between drug and disease nodes using homogeneous graphs in Euclidean space. However, linear transformations and distance metrics in Euclidean space oversimplify the similarity relationships between nodes. This simplification can lead to the model failing to accurately capture the complex association patterns between nodes in the homogeneous graph, hindering the learning of high-quality similarity features between drug and disease nodes, ultimately impacting the accuracy of drug-disease association prediction.
[0007] 2. Existing drug-disease association prediction models typically learn heterogeneous features of drug and disease nodes through heterogeneous graphs in Euclidean space. However, heterogeneous drug and disease graphs often contain hierarchical structures, and the volume of the graph increases exponentially with radius, while the volume of a sphere in Euclidean space increases polynomially with radius. This difference makes Euclidean space prone to modeling distortion. Specifically, the geometric properties of Euclidean space do not match the hierarchical, tree-like structure of heterogeneous graphs, resulting in the model's inability to accurately capture the semantic associations and topological relationships between nodes. These semantic associations and topological relationships are crucial for understanding the complex relationships between drug and disease nodes. Therefore, the limitations of Euclidean space affect the model's ability to predict drug-disease associations.
[0008] 3. Limited feature representation capabilities. Although deep learning methods have improved feature extraction through their powerful representation learning capabilities, many methods still fail to fully integrate the similar features and heterogeneous characteristics of drug nodes and disease nodes, resulting in the model being unable to generate comprehensive and high-quality feature representations.
[0009] 4. The importance of negative samples in drug-disease association prediction tasks has been widely recognized, but selecting the most representative negative samples remains a challenge. Existing methods typically train by randomly sampling negative samples, but this approach can overlook important negative samples or introduce irrelevant noise samples, thus affecting model performance.
[0010] Therefore, a drug-disease association prediction method based on a hyperbolic graph feature learning network is invented. Summary of the Invention
[0011] To solve the above technical problems, according to one aspect of the present invention, the present invention provides the following technical solutions:
[0012] A method for predicting drug-disease associations based on a hyperbolic graph feature learning network includes the following specific steps:
[0013] S1: Obtain drug-disease association data from the database to construct the drug similarity matrix DR and the disease similarity matrix DS. After extracting their features, map them into the hyperbolic space to obtain the hyperbolic drug initial feature matrix ZR H , hyperbolic disease initial characteristic matrix ZD H , used to construct the drug-disease adjacency matrix A;
[0014] S2: To ZR H and ZD H Apply the hyperbolic graph feature reconstructor to obtain the hyperbolic drug similarity feature matrix SR H and hyperbolic disease similarity matrix SD H ;
[0015] S3: Construct a heterogeneous graph AG based on A, and use GCN and hyperbolic heterogeneous change graph converter to transform ZR based on AG. H and ZD H The hyperbolic drug heterogeneous feature matrix HR is obtained by training H and the hyperbolic disease heterogeneous feature matrix HD H ;
[0016] S4: Splicing SR H and HR H , splicing SD H and HD H , use MLP to get the hyperbolic drug feature matrix FR H and the hyperbolic disease characteristic matrix FD H ;
[0017] S5: Based on FR H and FD H , using the positive and negative fusion hard case sampling strategy to construct the hyperbolic hard case drug feature matrix FR' H and the hyperbolic hard-to-treat disease feature matrix FD'H ;
[0018] S6: Utilize FR' H and FD' H , output the prediction results through MLP.
[0019] As a preferred solution of the method for predicting the association between drugs and diseases based on a hyperbolic graph feature learning network according to the present invention, the specific steps of S1 are as follows:
[0020] S11: Construct the drug similarity matrix DR and the disease similarity matrix DS, and use them through MLP to obtain the drug initial feature matrix ZR and the disease initial feature matrix ZD;
[0021] S12: Mapping the drug initial feature matrix ZR and the disease initial feature matrix ZD to the hyperbolic space;
[0022] S13: Constructing a drug-disease adjacency matrix
[0023] As a preferred solution of the method for predicting the association between drugs and diseases based on a hyperbolic graph feature learning network according to the present invention, the specific steps of S11 are as follows:
[0024] S111: Obtain drug-disease association data from the database and construct a drug similarity matrix Disease similarity matrix Where nr represents the number of drug nodes, and nd represents the number of disease nodes;
[0025] S112: Use MLP to extract features from DR and DS to obtain the initial drug feature matrix ZR and the initial disease feature matrix ZD. The formula is as follows:
[0026] ZR=MLP(DR)
[0027] ZD=MLP(DS)
[0028] in, represents the initial characteristic matrix of the drug, d0 is its characteristic dimension, Represents the initial feature vector of the drug in row i of ZR, i.e., the drug node dr i The initial feature vector of the drug; represents the initial disease feature matrix; Represents the initial feature vector of the disease in row i of ZD, i.e., the disease node ds i The initial feature vector of the disease;
[0029] The specific steps of S12 are as follows:
[0030] S121: Using the Lorentz model, let K = 1, then Expressed as The origin of express
[0031] S122: For ZR i , mapping it into the hyperbolic space, the formula is as follows:
[0032]
[0033] in, Represents the drug node dr i The initial eigenvector of the hyperbolic drug, exp o (·) is used to map variables to hyperbolic space (0, ZR i )satisfy represents the Minkowski inner product; ZR i Represents the drug node dr i The initial feature vector of the drug; ||.||2 represents the L2 norm;
[0034] S123: By any drug node dr i The initial drug feature vector ZR i By performing the above operations, we can obtain the initial characteristic matrix of hyperbolic drugs
[0035] S124: Repeat the above steps to obtain the initial feature matrix of the hyperbolic disease
[0036] The specific steps of S13 are as follows:
[0037] Obtain drug-disease association data and construct the adjacency matrix A. The specific construction method is as follows:
[0038]
[0039] As a preferred solution of the method for predicting the association between drugs and diseases based on a hyperbolic graph feature learning network according to the present invention, the specific steps of S2 are as follows:
[0040] S21: Use hyperbolic feature transformation to extract intermediate feature vectors of drug nodes and disease nodes;
[0041] S22: Calculate the hyperbolic homogeneous fusion attention and use the hyperbolic homogeneous fusion attention to extract the hyperbolic similarity features of drug and disease nodes.
[0042] As a preferred embodiment of the method for predicting the association between drugs and diseases based on a hyperbolic graph feature learning network according to the present invention, the specific steps of S21 are as follows:
[0043] S211: For nr drug nodes in the drug-disease association data, the hyperbolic drug initial feature matrix ZR is known. H ; Indicates ZR H The initial eigenvector of the hyperbolic drug in the i-th row;
[0044] At level 0, the hyperbolic drug similarity feature matrix is as follows:
[0045] SR 0,H =ZR H
[0046] but Indicates SR 0,H The hyperbolic drug similarity feature vector of the i-th row;
[0047] S212: For any drug node dr in the drug-disease association data i In the hyperbolic graph feature reconstructor of the first layer, the hyperbolic feature transformation is first performed to map the hyperbolic space of the previous layer to the hyperbolic space of this layer. The hyperbolic feature transformation formula of the first layer is as follows:
[0048]
[0049] in, Represented as drug node dr i After the hyperbolic characteristic transformation, the The intermediate eigenvector of ; is the trainable parameter of layer l, d l is the dimension of the lth layer; It is the drug node dr i Hyperbolic drug similarity feature vector at level l-1; b l The first layer is located at The Euclidean vector of is a trainable parameter; Represents the drug node dr i In the next layer l is located in the hyperbolic space Intermediate variables in exp o The role of (·) is to map the variables to the hyperbolic space of the next layer l log o The function of (·) is to map the variable to the tangent space of the origin o The effect is point As the center, map the variables to the hyperbolic space middle; b l The tangent space from the origin o Parallel transport to point The tangent space
[0050] The specific steps of S22 are as follows:
[0051] S221, calculating hyperbolic homogeneous fusion attention, where the hyperbolic homogeneous fusion attention consists of node feature attention, hyperbolic distance attention, and structural attention;
[0052] S2211: Calculate the drug node dr in the lth layer i For drug node dr j Node feature attention The specific formula is:
[0053]
[0054] Among them, LeakReLU is the activation function; is a trainable parameter; log o The function of (·) is to map the variable to the tangent space of the origin o express and Perform matrix multiplication.
[0055] Then normalize to get the final node feature attention The formula is as follows:
[0056]
[0057] in, Represents the drug node dr i The first-order neighbor set of , including itself; when the drug node dr i and drug node dr j The higher the similarity of node features in layer l, The bigger;
[0058] S2212: Calculate the drug node dr in the lth layer i For drug node dr j Hyperbolic distance attention The specific formula is:
[0059]
[0060] Among them, δ and η are hyperparameters used to adjust the hyperbolic distance; Used for calculation and The hyperbolic distance between represents the square of the hyperbolic distance;
[0061] Then normalize to get the final hyperbolic distance attention The specific formula is:
[0062]
[0063] in, Represents the drug node dr i The first-order neighbor set of , including itself; when the drug node dr i and drug node dr j The smaller the hyperbolic distance at level l, The bigger;
[0064] S2213: Calculate drug node dr i For drug node dr j The structural attention of , the specific calculation process is as follows:
[0065] Known drug similarity matrix DR, drug node dr i The degree of is defined here as:
[0066]
[0067] The formula means to sum each column j of the i-th row of DR, that is, the drug node dr i For each drug node dr j Sum the similarities;
[0068] Drug Node dr i For drug node dr j Structural Attention The calculation formula is as follows:
[0069]
[0070] in,
[0071] Then normalize to get the final structural attention The formula is as follows:
[0072]
[0073] in, Represents the drug node dr i The first-order neighbor set of , including itself; when the drug node dr i and drug node dr j The smaller the degree of The bigger;
[0074] S2214: At layer l, attention is paid to node features Hyperbolic distance attention and structural attention The fusion obtains the hyperbolic homogeneous fusion attention, the formula is as follows:
[0075]
[0076] where w NF 、w DS and w ST It is a trainable parameter and is initialized to 1;
[0077] S222: Extracting hyperbolic similarity features of drug and disease nodes using hyperbolic homogeneous fusion attention;
[0078] S2221: Yes Perform hyperbolic encoding to obtain the drug node dr i Hyperbolic similarity features at level l The formula for hyperbolic encoding is as follows:
[0079]
[0080] in, Its function is to map variables to hyperbolic space In the example, σ is the activation function; N i Represents the drug node dr i The first-order neighbor set of , including itself; The function is to map variables to The tangent space middle;
[0081] S2222: For any drug node dr i Perform the above operations and iterate through h layers to finally obtain the hyperbolic drug similarity feature matrix use Indicates SR h,H ;
[0082] S2223: Repeat the above steps to obtain the hyperbolic disease similarity feature matrix
[0083] As a preferred solution of the method for predicting the association between drugs and diseases based on a hyperbolic graph feature learning network according to the present invention, the specific steps of S3 are as follows:
[0084] S31: Use GCN to obtain hyperbolic disease heterogeneous initial feature matrix and hyperbolic drug heterogeneous initial feature matrix;
[0085] S32: Calculate the hyperbolic heterogeneous changing attention and use the hyperbolic heterogeneous changing attention to learn the heterogeneous features of drug nodes and disease nodes.
[0086] As a preferred embodiment of the method for predicting the association between drugs and diseases based on a hyperbolic graph feature learning network according to the present invention, the specific steps of S31 are as follows:
[0087] S311: Construct a heterogeneous graph AG based on the adjacency matrix A as follows:
[0088]
[0089] in,
[0090] S312: In order to reflect the number of layers, let MR 0,H Expressed as ZR H , MD 0,H Indicated as ZD H ; In order to use the GCN formula, MR 0,H ,MD 0,H After splicing, we get:
[0091]
[0092] in, E 0 It is the 0th layer by MR 0,H ,MD 0,H The concatenated feature matrix;
[0093] S313: At layer l, GCN is used to obtain E based on the heterogeneous graph AG. l , the formula is as follows:
[0094]
[0095] AG'=(AG+I)
[0096]
[0097] in, It is E l The hyperbolic heterogeneous initial eigenvector in the i-th row of layer l; σ is the activation function; exp o (·) is used to map variables to hyperbolic space It is the drug node dr i or disease node ds i The set of first-order neighbors of , including itself; is the processed adjacency matrix, yes The value of row i and column j; I is the identity matrix; is a trainable parameter; log o (·) is used to map the variable to the tangent space of the origin o The D matrix is used to calculate
[0098] S314: After h layers of iteration, we can get:
[0099]
[0100] That is, from E h The hyperbolic drug heterogeneous initial characteristic matrix is obtained from and hyperbolic disease heterogeneous initial feature matrix
[0101] The specific steps of S32 are as follows:
[0102] Given the hyperbolic drug heterogeneous initial feature matrix and the hyperbolic disease heterogeneous initial feature matrix, namely MR h,H and MD h,H ;
[0103] At level 0, the hyperbolic heterogeneous drug feature matrix and the hyperbolic heterogeneous disease feature matrix are:
[0104] HR 0,H =MR h,H
[0105] HD 0,H =MD h,H
[0106] The hyperbolic heterogeneous change graph converter has three inputs: the query matrix Q 1 and Q 2 , key matrix K 1 and K 2 And the value matrix V; if the hyperbolic drug heterogeneous features are extracted, the query matrix is obtained by HR 0,H The key matrix and value matrix are obtained through HD 0,H If we extract the hyperbolic disease heterogeneous features, the situation is just the opposite; at the lth layer, the outputs are HR l,H and HD l,H ;
[0107] At level l, we get HR l,H The process is as follows:
[0108]
[0109] Among them, HR l-1,H is the hyperbolic drug heterogeneous feature matrix at level l-1; exp o (·) is used to map variables to hyperbolic space log o (·) is used to map the variable to the tangent space of the origin o express and Perform matrix multiplication, Similarly;
[0110]
[0111] Among them, HD l-1,H is the hyperbolic disease heterogeneous feature matrix at level l-1; The meaning is the same as above;
[0112] Define the drug node dr in layer l i The update equation is as follows:
[0113]
[0114] in, Represents the drug node dr at layer l i Hyperbolic drug heterogeneous knowledge feature vector; V j is the jth row of V; represents the Lorentz norm; It's HR l,H Hyperbolic drug heterogeneity feature vector of row i; Norm is LayerNorm or BatchNorm; exp o (·) is used to map variables to hyperbolic space log o (·) is used to map the variable to the tangent space of the origin o
[0115] RDIF ij Defined as:
[0116]
[0117] RDIF ij is the hyperbolic attention score, λ l Defined as:
[0118]
[0119] Among them, λ l is a scalar at level l, is the learnable parameter of layer l, λ init is a hyperparameter;
[0120] After h layers of iteration, the drug node dr i The output hyperbolic drug heterogeneity feature vector is Then we get the hyperbolic drug heterogeneous feature matrix Use HR H HR h,H ;
[0121] Similarly, the hyperbolic disease heterogeneous feature matrix HD can also be obtained at the hth layer H ;
[0122] As a preferred solution of the method for predicting the association between drugs and diseases based on a hyperbolic graph feature learning network according to the present invention, the specific steps of S4 are as follows:
[0123] Calculate the hyperbolic drug and disease feature matrix as follows:
[0124] FR H =exp o (MLP(log o (SR H )||log o (HR H )))
[0125] FD H =exp o (MLP(log o (SD H )||log o (HD H )))
[0126] Among them, exp o (·) is used to map variables to hyperbolic space log o (·) is used to map the variable to the tangent space of the origin o The role of MLP is to transform the feature vectors of hyperbolic drugs and diseases from Dimensionality reduction to
[0127] As a preferred solution of the method for predicting the association between drugs and diseases based on a hyperbolic graph feature learning network according to the present invention, the specific steps of S5 are as follows:
[0128] S51: Construct drug negative samples and disease negative samples;
[0129] S52: Use the positive-negative fusion hard sample sampling strategy to generate hard drug negative samples and hard disease negative samples and obtain the hyperbolic hard drug feature matrix and the hyperbolic hard disease feature matrix;
[0130] The specific steps of S51 are as follows:
[0131] 511: First, the hyperbolic drug feature matrix and the hyperbolic disease signature matrix Mapped into drug feature matrix XR and disease feature matrix XD, as follows:
[0132] XR=log o (FR H )
[0133] XD=log o (FD H )
[0134] in, log o The function is to map the variable to the tangent space of the origin o middle;
[0135] S512: Obtain drug-disease association data and generate a heterogeneous graph G, which includes drug-disease relationships, drug-drug relationships, and disease-disease relationships; find all positive sample pairs and negative sample pairs from the heterogeneous graph G. First, ensure that the number of positive sample pairs and negative sample pairs is the same. The strategy adopted is: randomly select a part of negative sample pairs from the heterogeneous graph G to ensure that the number is the same as that of positive sample pairs, and vice versa; finally, select drug negative samples and disease negative samples from the selected negative sample pairs; where the drug negative sample feature matrix is The disease negative sample feature matrix is N neg is the number of negative sample pairs, that is, the number of drug negative samples and disease negative samples;
[0136] The specific steps of S52 are as follows:
[0137] S521: Construct the feature matrix of hard negative candidate samples;
[0138] For any node in the heterogeneous graph G, the set consisting of all first-order and second-order neighbor nodes is called the neighbor set of the node;
[0139] For all drug negative samples, find all its first-order and second-order neighbors, that is, the neighbor set;
[0140] For any drug negative sample r, the operation of constructing the feature matrix of hard candidate negative samples is as follows:
[0141] First, find all drug positive samples in its neighbor set according to the heterogeneous graph G. The local neighbor drug positive sample feature matrix composed of its drug positive samples is: XR pl , then fuse the feature vectors of all drug positive samples to obtain the drug fusion positive sample feature vector xpr , the formula is as follows:
[0142]
[0143] in It's XR pl The drug positive sample feature vector of the mr-th row; mr represents the number of drug positive samples in the neighbor set of drug negative sample r;
[0144] Similarly, the feature vectors of all disease positive samples in its neighbor set are fused to obtain the disease fusion positive sample feature vector x pd , the formula is as follows:
[0145]
[0146] in Yes XD pl The disease positive sample feature vector of the md-th row; md represents the number of disease positive samples in the neighbor set of the drug negative sample r;
[0147] Then find the disease negative samples of the second-order neighbors in the drug negative sample r neighbor set, and the second-order neighbor disease negative sample feature matrix composed of them is: XD nl ; According to x pr and x pd The probability distribution of disease negative samples is calculated for the disease negative samples of the second-order neighbors in the drug negative sample r neighbor set. The formula is as follows:
[0148]
[0149] in, XD nl The qth row feature vector of ; μ is a hyperparameter used to balance the effects of drugs and diseases; N represents the number of disease negative samples of the second-order neighbors in the neighbor set of drug negative samples r; The meaning of this formula is the distance (x pr ,x pd )The closer the disease negative sample is, the greater the probability of being selected as the disease negative sample;
[0150] Then according to the probability distribution of disease negative samples Select the disease negative samples of the second-order neighbors in the neighbor set of drug negative samples r corresponding to the M largest probability values, and construct a candidate disease negative sample feature matrix of size M M is a hyperparameter;
[0151] Next, the disease fusion positive sample feature vector x pd The information is injected into the candidate disease negative sample feature matrix T to construct the difficult candidate disease negative sample feature matrix T'. The formula is as follows:
[0152] T' i =αx pd +(1-α)T i ,α∈(0,1)
[0153] Among them, T' i is the negative sample feature vector of the hard disease in the i-th row of T'; α is a hyperparameter; T i is the negative sample feature vector of candidate disease in row i of T;
[0154] S522: Obtaining a hyperbolic difficult-case drug feature matrix and a hyperbolic difficult-case disease feature matrix;
[0155] For each row of the difficult candidate disease negative sample feature matrix T', the difficult disease negative sample feature vector T' j The drug negative sample feature vector of the drug negative sample r Dot product, the formula is as follows:
[0156]
[0157] in, It is the feature vector of the difficult disease negative sample with the largest dot product selected from T';
[0158] Since there are N negative samples of drugs neg So we can generate N neg The hard-to-detect disease negative sample feature vectors constitute the hard-to-detect disease negative sample feature matrix
[0159] Similarly, for disease negative samples, N neg difficult drug negative sample feature vectors, forming a difficult drug negative sample feature matrix
[0160] Finally based on XR' n and XD' n Update XR and XD to obtain the difficult drug feature matrix XR' and the difficult disease feature matrix XD', where XR' is the same size as XR, and XD' is the same size as XD;
[0161] Then map it back to the hyperbolic space to get the hyperbolic difficult drug feature matrix FR' H and the hyperbolic hard-to-treat disease feature matrix FD' H ,as follows:
[0162] FR' H =exp o (XR')
[0163] FD' H =exp o (XD')
[0164] where exp o Its function is to map variables to hyperbolic space middle.
[0165] As a preferred solution of the method for predicting the association between drugs and diseases based on a hyperbolic graph feature learning network according to the present invention, the specific steps of S6 are as follows:
[0166] S61: Calculate FR' H and FD' H Hyperbolic distance ij ;
[0167] Compute hyperbolic hard-to-treat drug feature vectors and hyperbolic hard-to-treat disease feature vectors Hyperbolic distance ij , the calculation formula is as follows:
[0168]
[0169] in, Distance ij The smaller it is, the higher the possibility of drug-disease association;
[0170] S62: Using Distance, determine the predicted drug-disease association probability output through MLP;
[0171] output = MLP(Distance)
[0172] in,
[0173] Compared with existing technologies:
[0174] The present invention introduces hyperbolic homogeneous fusion attention through a hyperbolic graph feature reconstructor, and learns compact and representative similar features of drug nodes and disease nodes in a homogeneous graph based on hyperbolic space, thereby more accurately capturing the similarity between drug nodes and drug nodes, and disease nodes and disease nodes, and avoiding the oversimplification problem when the drug and disease association prediction model learns similar features between drug nodes and drug nodes, and disease nodes and disease nodes in Euclidean space; the hyperbolic heterogeneous change graph converter introduces hyperbolic heterogeneous change attention, thereby more accurately capturing the semantic association and topological relationship between drug nodes and disease nodes on the heterogeneous graph, and semantic association and topological relationship are important components of heterogeneous features, so compact and representative heterogeneous features can be generated, avoiding the distortion problem caused by the drug and disease association prediction model when learning heterogeneous features of drug nodes and disease nodes in a heterogeneous graph, and also avoiding the attention noise problem; the hyperbolic collaborative representation learning strategy fuses the similar features learned by the hyperbolic graph feature reconstructor and the heterogeneous features learned by the hyperbolic heterogeneous change graph converter through MLP. Through this strategy, the model can generate comprehensive and high-quality features, thereby enabling a more comprehensive and in-depth exploration of the relationship between drugs and diseases, significantly improving the generalization ability of the model, and avoiding problems such as the model's inability to generate comprehensive and high-quality features; the positive-negative fusion difficult example sampling strategy synthesizes the most informative negative samples, significantly improving the representativeness and information content of negative samples, enabling the model to better distinguish the boundaries between positive and negative samples, avoiding problems such as ignoring important negative samples or introducing irrelevant noise negative samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0175] Figure 1 This is a main structure diagram of the hyperbolic graph feature learning network of the present invention;
[0176] Figure 2 It is a schematic diagram of the process of the present invention;
[0177] Figure 3 Schematic diagram of the structure of the hyperbolic graph feature reconstructor of the present invention;
[0178] Figure 4 This is a schematic structural diagram of the hyperbolic heterogeneous change graph converter of the present invention;
[0179] Figure 5 Schematic diagram of the positive and negative fusion difficult example sampling strategy of the present invention. DETAILED DESCRIPTION
[0180] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0181] The present invention provides a method for predicting the association between drugs and diseases based on a hyperbolic graph feature learning network. Figure 1-Figure 5, including the following specific steps:
[0182] S1: Obtain drug-disease association data from the database to construct the drug similarity matrix DR and the disease similarity matrix DS. After extracting their features, map them into the hyperbolic space to obtain the hyperbolic drug initial feature matrix ZR H , hyperbolic disease initial characteristic matrix ZD H , used to construct the drug-disease adjacency matrix A;
[0183] The specific steps of S1 are as follows:
[0184] S11: Construct the drug similarity matrix DR and the disease similarity matrix DS, and use them through MLP to obtain the drug initial feature matrix ZR and the disease initial feature matrix ZD;
[0185] The specific steps of S11 are as follows:
[0186] S111: Obtain drug-disease association data from the database and construct a drug similarity matrix Disease similarity matrix Where nr represents the number of drug nodes, and nd represents the number of disease nodes;
[0187] S112: Use MLP to extract features from DR and DS to obtain the initial drug feature matrix ZR and the initial disease feature matrix ZD. The formula is as follows:
[0188] ZR=MLP(DR)
[0189] ZD=MLP(DS)
[0190] in, represents the initial characteristic matrix of the drug, d0 is its characteristic dimension, Represents the initial feature vector of the drug in row i of ZR, i.e., the drug node dr i The initial feature vector of the drug; represents the initial disease feature matrix; Represents the initial feature vector of the disease in row i of ZD, i.e., the disease node ds i The initial feature vector of the disease;
[0191] Among them, the Multilayer Perceptron (MLP) is a basic neural network architecture that is widely used in data analysis and pattern recognition tasks. It can capture the nonlinear relationship between input and output through multiple levels of node connections, thereby realizing complex data modeling and learning functions.
[0192] A homogeneous graph is a graph data structure that contains only a single type of nodes and a single type of edges. The drug similarity matrix DR corresponds to the drug homogeneous graph, and the disease similarity matrix DS corresponds to the disease homogeneous graph.
[0193] Example: The C-dataset dataset is used in the experiments of the present invention. The dataset consists of 663 drug nodes, 409 disease nodes, and 993 protein nodes, as well as 2532 verified drug-disease association pairs, 3773 drug-protein association pairs, and disease-protein association pairs. The molecular fingerprint and GIP similarity are used to calculate the similarity between drugs. If the molecular fingerprint similarity between two drugs is 0, the GIP similarity is used as the drug similarity matrix, otherwise the average value is used as the drug similarity matrix. The phenotypic and GIP similarity are used to calculate the similarity between diseases. If the phenotypic similarity between two diseases is 0, the GIP similarity is used as the disease similarity matrix, otherwise the average value is used as the disease similarity matrix. In the C-dataset dataset, the drug similarity matrix DR is 663×663 in size, and the disease similarity matrix DS is 409×409 in size.
[0194] S12: Mapping the drug initial feature matrix ZR and the disease initial feature matrix ZD to the hyperbolic space;
[0195] The specific steps of S12 are as follows:
[0196] S121: Using the Lorentz model, let K = 1, then o = {1, 0, ..., 0} ∈ Expressed as The origin of express
[0197] There are many equivalent hyperbolic models in hyperbolic space, such as the Poincaré ball model, the Lorentz model, and the Klein model. These models exhibit different characteristics. For the Lorentz model, the constant negative curvature of the hyperbolic space is -1 / K (K>0). Let the origin be Hyperbolic space The origin of , where d represents the dimension of the hyperbolic space, and K represents the constant negative curvature of the hyperbolic space, which is -1 / K (K>0);
[0198] S122: For ZR i , mapping it into the hyperbolic space, the formula is as follows:
[0199]
[0200] in, Represents the drug node dr i The initial eigenvector of the hyperbolic drug, exp o (·) is used to map variables to hyperbolic space (0, ZR i )satisfy represents the Minkowski inner product; ZR i Represents the drug node dr i The initial feature vector of the drug; ||.||2 represents the L2 norm;
[0201] Among them, this is the formula for mapping eigenvectors to hyperbolic space in the Lorentz model when the constant negative curvature is -1; the Minkowski inner product is a generalized inner product definition, often used to measure similarity or distance in vector space. It is a generalized form of the Euclidean inner product (dot product). In the Lorentz model, it is defined as:
[0202] S123: By any drug node dr i The initial drug feature vector ZR i By performing the above operations, we can obtain the initial characteristic matrix of hyperbolic drugs
[0203] S124: Repeat the above steps to obtain the initial feature matrix of the hyperbolic disease
[0204] S13: Constructing a drug-disease adjacency matrix
[0205] The specific steps of S13 are as follows:
[0206] Obtain drug-disease association data and construct the adjacency matrix A. The specific construction method is as follows:
[0207]
[0208] Example: In the experiment of the present invention, a C-dataset dataset was obtained, which consists of 663 drug nodes, 409 disease nodes, 993 protein nodes, and 2532 verified drug-disease association pairs, 3773 drug-protein association pairs, and disease-protein association pairs; an adjacency matrix A was constructed based on the 2532 drug-disease association pairs;
[0209] S2: To ZR H and ZD HApply the hyperbolic graph feature reconstructor to obtain the hyperbolic drug similarity feature matrix SR H and hyperbolic disease similarity matrix SD H ;
[0210] The present invention designs a hyperbolic graph feature reconstructor (HGFR), which is a hyperbolic graph feature reconstructor based on GAT design. It innovatively designs hyperbolic homogeneous fusion attention (HHFA) for extracting similar features of drug nodes and disease nodes in hyperbolic space through homogeneous graphs, that is, hyperbolic similarity features. The steps are to use hyperbolic feature transformation to extract intermediate feature vectors of drug nodes and disease nodes; calculate hyperbolic homogeneous fusion attention, and use hyperbolic homogeneous fusion attention to extract hyperbolic similarity features of drug and disease nodes. The structure of the hyperbolic graph feature reconstructor is as follows: Figure 3 As shown:
[0211] The specific steps of S2 are as follows:
[0212] S21: Use hyperbolic feature transformation to extract intermediate feature vectors of drug nodes and disease nodes;
[0213] The specific steps of S21 are as follows:
[0214] S211: For nr drug nodes in the drug-disease association data, the hyperbolic drug initial feature matrix ZR is known. H ; Indicates ZR H The initial eigenvector of the hyperbolic drug in the i-th row;
[0215] At level 0, the hyperbolic drug similarity feature matrix is as follows:
[0216] SR 0,H =ZR H
[0217] but Indicates SR 0,H The hyperbolic drug similarity feature vector of the i-th row;
[0218] S212: For any drug node dr in the drug-disease association data i In the hyperbolic graph feature reconstructor of the first layer, the hyperbolic feature transformation is first performed to map the hyperbolic space of the previous layer to the hyperbolic space of this layer. The hyperbolic feature transformation formula of the first layer is as follows:
[0219]
[0220] Among them, the feature transformation formula of the hyperbolic space is specifically used to map the feature vector to the next layer of hyperbolic space, which is named hyperbolic feature transformation here; the logarithmic mapping log o The function of (·) is to map the variable to the tangent space, that is, the Euclidean space; the exponential mapping exp o The role of (·) is to map variables into hyperbolic space;
[0221] in, Represented as drug node dr i After the hyperbolic characteristic transformation, the The intermediate eigenvector of ; is the trainable parameter of layer l, d l is the dimension of the lth layer; It is the drug node dr i Hyperbolic drug similarity feature vector at level l-1; b l The first layer is located at The Euclidean vector of is a trainable parameter; Represents the drug node dr i In the next layer l is located in the hyperbolic space Intermediate variables in exp o The role of (·) is to map the variables to the hyperbolic space of the next layer l log o The function of (·) is to map the variable to the tangent space of the origin o The effect is point As the center, map the variables to the hyperbolic space middle; Will The tangent space from the origin o Parallel transport to point The tangent space
[0222] S22: Calculate the hyperbolic homogeneous fusion attention and use the hyperbolic homogeneous fusion attention to extract the hyperbolic similarity features of drug and disease nodes;
[0223] The specific steps of S22 are as follows:
[0224] S221, calculating hyperbolic homogeneous fusion attention, where the hyperbolic homogeneous fusion attention consists of node feature attention, hyperbolic distance attention, and structural attention;
[0225] S2211: Calculate the drug node dr in the lth layer i For drug node dr j Node feature attention The specific formula is:
[0226]
[0227] Among them, LeakReLU is the activation function; It is a trainable parameter; the role of log0(·) is to map the variable to the tangent space of the origin o express and Perform matrix multiplication.
[0228] Then normalize to get the final node feature attention The formula is as follows:
[0229]
[0230] in, Represents the drug node dr i The first-order neighbor set of , including itself; when the drug node dr i and drug node dr j The higher the similarity of node features in layer l, The bigger;
[0231] S2212: Calculate the drug node dr in the lth layer i For drug node dr j Hyperbolic distance attention The specific formula is:
[0232]
[0233] Among them, δ and η are hyperparameters used to adjust the hyperbolic distance; Used for calculation and The hyperbolic distance between represents the square of the hyperbolic distance;
[0234] Then normalize to get the final hyperbolic distance attention The specific formula is:
[0235]
[0236] in, Represents the drug node dr i The first-order neighbor set of , including itself; when the drug node dr i and drug node dr j The smaller the hyperbolic distance at level l, The bigger;
[0237] S2213: Calculate drug node dr i For drug node dr jThe structural attention of , the specific calculation process is as follows:
[0238] Known drug similarity matrix DR, drug node dr i The degree of is defined here as:
[0239]
[0240] The formula means to sum each column j of the i-th row of DR, that is, the drug node dr i For each drug node dr j Sum the similarities;
[0241] Drug Node dr i For drug node dr j Structural Attention The calculation formula is as follows:
[0242]
[0243] in,
[0244] Then normalize to get the final structural attention The formula is as follows:
[0245]
[0246] in, Represents the drug node dr i The first-order neighbor set of , including itself; when the drug node dr i and drug node dr j The smaller the degree of The bigger;
[0247] S2214: At layer l, attention is paid to node features Hyperbolic distance attention and structural attention The fusion obtains the hyperbolic homogeneous fusion attention, the formula is as follows:
[0248]
[0249] where w NF 、w DS and w ST It is a trainable parameter and is initialized to 1;
[0250] S222: Extracting hyperbolic similarity features of drug and disease nodes using hyperbolic homogeneous fusion attention;
[0251] S2221: Yes Perform hyperbolic encoding to obtain the drug node dri Hyperbolic similarity features at level l The formula for hyperbolic encoding is as follows:
[0252]
[0253] in, Its function is to map variables to hyperbolic space In the middle; σ is the activation function; Represents the drug node dr i The first-order neighbor set of , including itself; The function is to map variables to The tangent space middle;
[0254] S2222: For any drug node dr i Perform the above operations and iterate through h layers to finally obtain the hyperbolic drug similarity feature matrix use Indicates SR h,H ;
[0255] S2223: Repeat the above steps to obtain the hyperbolic disease similarity feature matrix
[0256] Example: The hyperbolic drug initial characteristic matrix ZR obtained by S1 H and the hyperbolic disease initial characteristic matrix ZD H As input, we get the hyperbolic drug similarity feature matrix and the hyperbolic disease similarity feature matrix at layer 0, ZR H Contains the hyperbolic drug initial feature vector of 663 drug nodes, whose feature dimension is d0, d0 is a hyperparameter; ZD H The hyperbolic disease initial feature vector contains 409 disease nodes, with a feature dimension of d0, where d0 is a hyperparameter. At level 0, the processing of drug nodes is as follows:
[0257] Step 1: For a certain drug node, use the hyperbolic feature transformation formula to map the hyperbolic drug similarity feature vector of the drug node to the next layer of hyperbolic space (i.e., the first layer) to obtain the intermediate feature vector of the drug node, whose feature dimension is d1;
[0258] Step 2: Based on the intermediate feature vector of the drug node, calculate the node feature attention of the first layer, the hyperbolic distance attention of the first layer, and the structural attention (structural attention is independent of the layer). Give different weights to these three types of attention (weights are trainable parameters), and finally perform weighted fusion to form a hyperbolic homogeneous fusion attention.
[0259] Step 3: Based on the intermediate feature vector of the drug node, the hyperbolic homogeneous fusion attention is used to update the hyperbolic drug similarity feature vector of the drug node to obtain the hyperbolic drug similarity feature vector of the first layer of the drug node;
[0260] Step 4: Repeat the first three steps for each row of the hyperbolic drug similarity feature vector of the hyperbolic drug similarity feature matrix of the 0th layer (that is, the hyperbolic drug similarity feature vector of each drug node) to obtain the hyperbolic drug similarity feature of the 1st layer; the remaining layers are similar to the above process. The above process is iterated h times to obtain the hyperbolic drug similarity feature matrix SR H , similarly, we can also get the hyperbolic disease similarity feature matrix SD H ;
[0261] S3: Construct a heterogeneous graph AG based on A, and use GCN and hyperbolic heterogeneous change graph converter to transform ZR based on AG. H and ZD H The hyperbolic drug heterogeneous feature matrix HR is obtained by training H and the hyperbolic disease heterogeneous feature matrix HD H ;
[0262] The present invention designs a hyperbolic heterogeneous dynamic graph transformer (HHDGT), which is a hyperbolic heterogeneous dynamic graph transformer designed based on Transformer. It innovatively introduces hyperbolic heterogeneous dynamic attention. It uses the hyperbolic drug heterogeneous initial feature moment and the hyperbolic disease heterogeneous initial feature matrix obtained by GCN training to train the hyperbolic drug heterogeneous feature matrix and the hyperbolic disease heterogeneous feature matrix. The process is to calculate the hyperbolic heterogeneous dynamic attention and use the hyperbolic heterogeneous dynamic attention to learn the heterogeneous features of drug nodes and disease nodes. The structure of the hyperbolic heterogeneous dynamic graph transformer is as follows: Figure 4 As shown:
[0263] The specific steps of S3 are as follows:
[0264] S31: Use GCN to obtain hyperbolic disease heterogeneous initial feature matrix and hyperbolic drug heterogeneous initial feature matrix;
[0265] The specific steps of S31 are as follows:
[0266] S311: Construct a heterogeneous graph AG based on the adjacency matrix A as follows:
[0267]
[0268] in,
[0269] S312: In order to reflect the number of layers, let MR 0,H Expressed as ZR H , MD 0,H Indicated as ZD H ; In order to use the GCN formula, MR 0,H ,MD 0,H After splicing, we get:
[0270]
[0271] in, E 0 It is the 0th layer by MR 0,H ,MD 0,H The concatenated feature matrix;
[0272] S313: At layer l, GCN is used to obtain E based on the heterogeneous graph AG. l , the formula is as follows:
[0273]
[0274] AG'=(AG+I)
[0275]
[0276]
[0277] GCN (Graph Convolutional Network) is a deep learning model for processing graph-structured data. It uses convolution operations to propagate information and extract features on graph structures, and is widely used in tasks such as node classification, graph classification, and link prediction. A heterogeneous graph is a graph-structured data containing multiple types of nodes (entities) and edges (relationships). Unlike homogeneous graphs (all nodes and edges are of the same type), nodes and edges in heterogeneous graphs have clear type distinctions, which can more naturally model the multivariate interactions in complex systems. In this heterogeneous graph, nodes of the same type are not connected.
[0278] in, It is E l The hyperbolic heterogeneous initial eigenvector in the i-th row of layer l; σ is the activation function; exp o (·) is used to map variables to hyperbolic space It is the drug node dr i or disease node ds i The set of first-order neighbors of , including itself; is the processed adjacency matrix, yes The value of row i and column j; I is the identity matrix; is a trainable parameter; log o (·) is used to map the variable to the tangent space of the origin o The D matrix is used to calculate
[0279] S314: After h layers of iteration, we can get:
[0280]
[0281] That is, from E h The hyperbolic drug heterogeneous initial characteristic matrix is obtained from and hyperbolic disease heterogeneous initial feature matrix
[0282] S32: Calculate hyperbolic heterogeneous changing attention and use it to learn heterogeneous features of drug nodes and disease nodes;
[0283] The specific steps of S32 are as follows:
[0284] Given the hyperbolic drug heterogeneous initial feature matrix and the hyperbolic disease heterogeneous initial feature matrix, namely MR h,H and MD h,H ;
[0285] At level 0, the hyperbolic heterogeneous drug feature matrix and the hyperbolic heterogeneous disease feature matrix are:
[0286] HR 0,H =MR h,H
[0287] HD 0,H =MD h,H
[0288] The hyperbolic heterogeneous change graph converter has three inputs: the query matrix Q 1 and Q 2 , key matrix K 1 and K 2 And the value matrix V; if the hyperbolic drug heterogeneous features are extracted, the query matrix is obtained by HR 0,H The key matrix and value matrix are obtained through HD 0,H If we extract the hyperbolic disease heterogeneous features, the situation is just the opposite; at the lth layer, the outputs are HR l,H and HD l,H ;
[0289] At level l, we get HR l,H The process is as follows:
[0290]
[0291] Among them, HRl-1,H is the hyperbolic drug heterogeneous feature matrix at level l-1; exp o (·) is used to map variables to hyperbolic space log o (·) is used to map the variable to the tangent space of the origin o express and Perform matrix multiplication, Similarly;
[0292]
[0293] Among them, HD l-1,H is the hyperbolic disease heterogeneous feature matrix at level l-1; The meaning is the same as above;
[0294] Define the drug node dr in layer l i The update equation is as follows:
[0295]
[0296] in, Represents the drug node dr at layer l i Hyperbolic drug heterogeneous knowledge feature vector; V j is the jth row of V; represents the Lorentz norm; It's HR l,H Hyperbolic drug heterogeneity feature vector of row i; Norm is LayerNorm or BatchNorm; exp o (·) is used to map variables to hyperbolic space log o (·) is used to map the variable to the tangent space of the origin o
[0297] BatchNorm is a deep learning technology proposed by Sergey Ioffe and Christian Szegedy in 2015 to solve the problem of gradient vanishing or gradient exploding in deep neural networks and accelerate the model training process. LayerNorm was proposed by Jimmy Lei Ba, Jamie Ryan Kiros and others in 2016; its main purpose is to help neural networks converge faster and more stably.
[0298] RDIF ij Defined as:
[0299]
[0300] RDIF ij is the hyperbolic attention score, λ l Defined as:
[0301]
[0302] Among them, λ l is a scalar at level l, is the learnable parameter of layer l, λ init is a hyperparameter;
[0303] After h layers of iteration, the drug node dr i The output hyperbolic drug heterogeneity feature vector is Then we get the hyperbolic drug heterogeneous feature matrix Use HR H HR h,H ;
[0304] Similarly, the hyperbolic disease heterogeneous feature matrix HD can also be obtained at the hth layer H ;
[0305] S4: Splicing SR H and HR H , splicing SD H and HD H , use MLP to get the hyperbolic drug feature matrix FR H and the hyperbolic disease characteristic matrix FD H ;
[0306] The specific steps of S4 are as follows:
[0307] Calculate the hyperbolic drug and disease feature matrix as follows:
[0308] FR H =exp o (MLP(log o (SR H )||log o (HR H )))
[0309] FD H =exp o (MLP(log o (SD H )||log o (HD H )))
[0310] Among them, exp o (·) is used to map variables to hyperbolic space log o (·) is used to map the variable to the tangent space of the origin o The role of MLP is to transform the feature vectors of hyperbolic drugs and diseases from Dimensionality reduction to
[0311] S5: Based on FR H and FD H , using the positive and negative fusion hard case sampling strategy to construct the hyperbolic hard case drug feature matrix FR' H and the hyperbolic hard-to-treat disease feature matrix FD' H ;
[0312] The present invention designs a positive-negative fusion difficult sampling strategy for generating difficult drug negative samples and difficult disease negative samples to improve the decision boundary of the model. The process includes constructing drug negative samples and disease negative samples; using the positive-negative fusion difficult sampling strategy to generate difficult drug negative samples and difficult disease negative samples and obtain hyperbolic difficult drug feature matrix and hyperbolic difficult disease feature matrix; the structure of the positive-negative fusion difficult sampling strategy is as follows Figure 5 As shown:
[0313] The specific steps of S5 are as follows:
[0314] S51: Construct drug negative samples and disease negative samples;
[0315] The specific steps of S51 are as follows:
[0316] S511: First, the hyperbolic drug feature matrix and the hyperbolic disease signature matrix Mapped into drug feature matrix XR and disease feature matrix XD, as follows:
[0317] XR=log o (FR H )
[0318] XD=log o (FD H )
[0319] in, log o The function is to map the variable to the tangent space of the origin o middle;
[0320] S512: Obtain drug-disease association data and generate a heterogeneous graph G, which includes drug-disease relationships, drug-drug relationships, and disease-disease relationships; find all positive sample pairs and negative sample pairs from the heterogeneous graph G. First, ensure that the number of positive sample pairs and negative sample pairs is the same. The strategy adopted is: randomly select a part of negative sample pairs from the heterogeneous graph G to ensure that the number is the same as that of positive sample pairs, and vice versa; finally, select drug negative samples and disease negative samples from the selected negative sample pairs; where the drug negative sample feature matrix is The disease negative sample feature matrix is N neg is the number of negative sample pairs, that is, the number of drug negative samples and disease negative samples;
[0321] A heterogeneous graph is a graph-structured data structure that contains multiple types of nodes (entities) and edges (relationships). Unlike homogeneous graphs (where all nodes and edges are of the same type), nodes and edges in a heterogeneous graph have clear type distinctions, enabling a more natural modeling of multivariate interactions in complex systems. In this heterogeneous graph G, nodes of the same type may be connected.
[0322] Example: In the experiment of the present invention, a C-dataset dataset was obtained, which consists of 663 drug nodes and 409 disease nodes. The adjacency matrix A can be obtained from S1. According to the adjacency matrix, 2532 drug-disease association pairs can be obtained, thereby obtaining the drug-disease relationship;
[0323] In S1, the drug similarity matrix and the disease similarity matrix are obtained. The drug-drug relationship can be obtained according to the drug similarity matrix, and the disease-disease relationship can be obtained according to the disease similarity matrix;
[0324] A heterogeneous graph G is constructed based on drug-drug relationships, disease-disease relationships, and drug-disease relationships. The heterogeneous graph G has 663 drug nodes and 409 disease nodes. There are 2532 verified drug-disease association pairs, so the number of positive sample pairs is 2532, and the number of negative sample pairs is 663×409-2532=268635.
[0325] Among the 268,635 negative sample pairs, 2,532 negative sample pairs were randomly selected to ensure that the number of positive sample pairs and negative sample pairs is the same, both 2,532;
[0326] S52: Use the positive-negative fusion hard sample sampling strategy to generate hard drug negative samples and hard disease negative samples and obtain the hyperbolic hard drug feature matrix and the hyperbolic hard disease feature matrix;
[0327] The specific steps of S52 are as follows:
[0328] S521: Construct the feature matrix of hard negative candidate samples;
[0329] For any node in the heterogeneous graph G, the set consisting of all first-order and second-order neighbor nodes is called the neighbor set of the node;
[0330] For all drug negative samples, find all its first-order and second-order neighbors, that is, the neighbor set;
[0331] For any drug negative sample r, the operation of constructing the feature matrix of hard candidate negative samples is as follows:
[0332] First, find all drug positive samples in its neighbor set according to the heterogeneous graph G. The local neighbor drug positive sample feature matrix composed of its drug positive samples is: XR pl , then fuse the feature vectors of all drug positive samples to obtain the drug fusion positive sample feature vector x pr , the formula is as follows:
[0333]
[0334] in It's XR pl The drug positive sample feature vector of the mr-th row; mr represents the number of drug positive samples in the neighbor set of drug negative sample r;
[0335] Similarly, the feature vectors of all disease positive samples in its neighbor set are fused to obtain the disease fusion positive sample feature vector x pd , the formula is as follows:
[0336]
[0337] in Yes XD pl The disease positive sample feature vector of the md-th row; md represents the number of disease positive samples in the neighbor set of the drug negative sample r;
[0338] Then find the disease negative samples of the second-order neighbors in the drug negative sample r neighbor set, and the second-order neighbor disease negative sample feature matrix composed of them is: XD nl ; According to x pr and x pd The probability distribution of disease negative samples is calculated for the disease negative samples of the second-order neighbors in the drug negative sample r neighbor set. The formula is as follows:
[0339]
[0340] in, XD nlThe qth row feature vector of ; μ is a hyperparameter used to balance the effects of drugs and diseases; N represents the number of disease negative samples of the second-order neighbors in the neighbor set of drug negative samples r; The meaning of this formula is the distance (x pr ,x pd )The closer the disease negative sample is, the greater the probability of being selected as the disease negative sample;
[0341] Then according to the probability distribution of disease negative samples Select the disease negative samples of the second-order neighbors in the neighbor set of drug negative samples r corresponding to the M largest probability values, and construct a candidate disease negative sample feature matrix of size M M is a hyperparameter;
[0342] Next, the disease fusion positive sample feature vector x pd The information is injected into the candidate disease negative sample feature matrix T to construct the difficult candidate disease negative sample feature matrix T'. The formula is as follows:
[0343] T' i =αx pd +(1-α)T i ,α∈(0,1)
[0344] Among them, T' i is the negative sample feature vector of the hard disease in the i-th row of T'; α is a hyperparameter; T i is the negative sample feature vector of candidate disease in row i of T;
[0345] S522: Obtaining a hyperbolic difficult-case drug feature matrix and a hyperbolic difficult-case disease feature matrix;
[0346] For each row of the difficult candidate disease negative sample feature matrix T', the difficult disease negative sample feature vector T' j The drug negative sample feature vector of the drug negative sample r Dot product, the formula is as follows:
[0347]
[0348] in, It is the feature vector of the difficult disease negative sample with the largest dot product selected from T';
[0349] Since there are N negative samples of drugs neg So we can generate N neg The hard-to-detect disease negative sample feature vectors constitute the hard-to-detect disease negative sample feature matrix
[0350] Similarly, for disease negative samples, N neg difficult drug negative sample feature vectors, forming a difficult drug negative sample feature matrix
[0351] Finally based on XR' n and XD' n Update XR and XD to obtain the difficult drug feature matrix XR' and the difficult disease feature matrix XD', where XR' is the same size as XR, and XD' is the same size as XD;
[0352] Then map it back to the hyperbolic space to get the hyperbolic difficult drug feature matrix FR' H and the hyperbolic hard-to-treat disease feature matrix FD' H ,as follows:
[0353] FR' H =exp o (XR')
[0354] FD' H =exp o (XD')
[0355] where exp o Its function is to map variables to hyperbolic space middle;
[0356] Example: 2532 drug negative samples and 2532 disease negative samples are screened out from the 2532 negative sample pairs screened out in Example S51, and then the positive-negative fusion hard sample sampling strategy is applied to the drug negative samples and the disease negative samples respectively; (drug, disease) node pairs that are not connected by any known association are called negative sample pairs. For example, there are 3 drug nodes and 3 disease nodes, and drug node 1 and disease node 1 are not connected. Then drug node 1 and disease node 1 are negative sample pairs, and drug node 1 and disease node 1 are called drug negative sample 1 and disease negative sample 1 respectively; negative samples are local. For example, although drug node 1 is a drug negative sample, it may still be connected to other disease nodes.
[0357] The process of using the positive-negative fusion hard sample sampling strategy for drug negative samples is as follows: for any drug negative sample r, the first step is to find all its first-order and second-order neighbors, that is, the neighbor set, and then fuse the feature vectors of all drug positive samples in the neighbor set (as long as there is a disease node connected to it, it is considered a drug positive sample) to obtain the drug fusion positive sample feature vector x pr , fuse the feature vectors of all disease positive samples in the neighbor set to obtain the disease fusion positive sample feature vector x pd (As long as there is a drug node connected to it, it is considered a positive disease sample);
[0358] Step 2: Find all the disease negative samples of the second-order neighbors (all the second-order neighbor disease nodes that are not connected to the drug negative sample node r), and then calculate these disease negative samples according to x pr and x pd Calculate the probability distribution of disease negative samples, and then select M disease negative samples with the largest probability values according to the probability distribution of disease negative samples;
[0359] Step 3: Change x pd Integrate into M disease negative samples to obtain M difficult disease negative samples, and integrate into x pd The feature vector of the hard disease negative sample after the calculation is called the hard disease negative sample feature vector;
[0360] Step 4: Perform dot products on the feature vectors of the M hard disease negative samples with the feature vector of the drug negative sample r, select the hard disease negative sample with the largest dot product value as the final result, and eliminate the remaining hard disease negative samples (the feature vectors of the eliminated hard disease negative samples will not be updated to the hard disease negative sample feature vector);
[0361] Step 5: Finally, 2532 drug negative samples are used to obtain 2532 difficult disease negative samples, and the difficult disease negative sample feature matrix XD' is obtained n , and then update the disease feature matrix XD to obtain the difficult disease feature matrix XD', and map it back to the hyperbolic space to obtain the hyperbolic difficult disease feature matrix FD' H If there are multiple drug negative samples that simultaneously select the same difficult disease negative sample (the difficult disease negative sample has multiple difficult disease negative sample feature vectors), then a difficult disease negative sample feature vector is randomly selected as the final difficult disease negative sample feature vector of the difficult disease negative sample;
[0362] Similarly, for disease negative samples, the positive and negative fusion difficult sample sampling strategy can also be used to obtain the difficult drug negative sample feature matrix XR' n , and then update the drug feature matrix XR to obtain the difficult drug feature matrix XR', and map it back to the hyperbolic space to obtain the hyperbolic difficult drug feature matrix FR' H ;
[0363] S6: Utilize FR' H and FD' H , output the prediction results through MLP;
[0364] The specific steps of S6 are as follows:
[0365] S61: Calculate FR' H and FD' H Hyperbolic distance ij ;
[0366] Compute hyperbolic hard-to-treat drug feature vectors and hyperbolic hard-to-treat disease feature vectors Hyperbolic distance ij , the calculation formula is as follows:
[0367]
[0368] in, Distance ij The smaller it is, the higher the possibility of drug-disease association;
[0369] S62: Using Distance, determine the predicted drug-disease association probability output through MLP;
[0370] output = MLP(Distance)
[0371] in,
[0372] Based on the above, the technical objectives of the present invention are:
[0373] 1. To address the oversimplification problem of drug-disease association prediction models when learning similarity features from homogeneous graphs in Euclidean space, we proposed a hyperbolic graph feature reconstructor. By introducing hyperbolic homogeneous fusion attention, the hyperbolic graph feature reconstructor learns compact and discriminative similarity features between drug and disease nodes from homogeneous graphs in hyperbolic space, thereby more accurately capturing the similarities between drug nodes and disease nodes.
[0374] 2. To address the problem of distortion in drug-disease association prediction models when learning heterogeneous features from heterogeneous graphs in Euclidean space, we proposed a hyperbolic heterogeneous change graph transformer. By designing hyperbolic heterogeneous change attention, the hyperbolic heterogeneous change graph transformer can learn the semantic associations and topological relationships of heterogeneous graphs in hyperbolic space. Semantic associations and topological relationships are important components of heterogeneous features. Therefore, it can generate compact and representative heterogeneous features, accurately modeling the complex relationships between drug nodes and disease nodes. This not only solves the distortion problem but also avoids the problem of attention noise.
[0375] 3. To address the model's inability to generate comprehensive and high-quality features due to limited feature representation capabilities, a hyperbolic collaborative representation learning strategy is proposed. The hyperbolic graph feature reconstructor learns similar features between drug and disease nodes in a homogeneous graph, while the hyperbolic heterogeneous change graph converter learns heterogeneous features between drug and disease nodes in a heterogeneous graph. Finally, the similar and heterogeneous features are fused. This strategy enables the model to generate comprehensive and high-quality features, enabling a more comprehensive and in-depth exploration of the associations between drug and disease nodes, significantly improving the model's generalization capabilities.
[0376] 4. To address the problem of selecting the most informative negative samples, we proposed a positive-negative fusion hard sampling strategy to synthesize difficult negative drug and disease samples. This strategy synthesizes the most informative negative samples by combining the probability distribution of negative samples and the fusion of positive samples. This strategy helps the model better distinguish between positive and negative sample pairs, thereby improving its predictive performance.
[0377] Although the present invention has been described above with reference to embodiments, various modifications may be made thereto and equivalent components may be substituted without departing from the scope of the present invention. In particular, as long as there are no structural conflicts, the various features of the embodiments disclosed herein may be combined with each other in any manner, and the omission of an exhaustive description of such combinations in this specification is solely for the sake of space and resource conservation. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A drug-disease association prediction method based on a hyperbolic graph feature learning network, characterized in that: The specific steps are as follows: S1: Obtain drug-disease association data from the database to construct the drug similarity matrix DR and the disease similarity matrix DS. After extracting their features, map them into the hyperbolic space to obtain the hyperbolic drug initial feature matrix ZR H , hyperbolic disease initial characteristic matrix ZD H , used to construct the drug-disease adjacency matrix A; S2: To ZR H and ZD H Apply the hyperbolic graph feature reconstructor to obtain the hyperbolic drug similarity feature matrix SR H and hyperbolic disease similarity matrix SD H ; S3: Construct a heterogeneous graph AG based on A, and use GCN and hyperbolic heterogeneous change graph converter to transform ZR based on AG. H and ZD H The hyperbolic drug heterogeneous feature matrix HR is obtained by training H and the hyperbolic disease heterogeneous feature matrix HD H ; S4: Splicing SR H and HR H , splicing SD H and HD H , use MLP to get the hyperbolic drug feature matrix FR H and the hyperbolic disease characteristic matrix FD H ; S5: Based on FR H and FD H , using the positive and negative fusion hard case sampling strategy to construct the hyperbolic hard case drug feature matrix FR' H and the hyperbolic hard-to-treat disease feature matrix FD' H ; S6: Utilize FR' H and FD' H , output the prediction results through MLP.
2. The method for predicting drug-disease association based on a hyperbolic graph feature learning network according to claim 1, characterized in that: The specific steps of S1 are as follows: S11: Construct the drug similarity matrix DR and the disease similarity matrix DS, and use them through MLP to obtain the drug initial feature matrix ZR and the disease initial feature matrix ZD; S12: Mapping the drug initial feature matrix ZR and the disease initial feature matrix ZD to the hyperbolic space; S13: Constructing a drug-disease adjacency matrix 3. The method for predicting drug-disease association based on a hyperbolic graph feature learning network according to claim 2, characterized in that: The specific steps of S11 are as follows: S111: Obtain drug-disease association data from the database and construct a drug similarity matrix Disease similarity matrix Where nr represents the number of drug nodes, and nd represents the number of disease nodes; S112: Use MLP to extract features from DR and DS to obtain the initial drug feature matrix ZR and the initial disease feature matrix ZD. The formula is as follows: ZR=MLP(DR) ZD=MLP(DS) in, represents the initial characteristic matrix of the drug, d0 is its characteristic dimension, Represents the initial feature vector of the drug in row i of ZR, i.e., the drug node dr i The initial feature vector of the drug; represents the initial disease feature matrix; Represents the initial feature vector of the disease in row i of ZD, i.e., the disease node ds i The initial feature vector of the disease; The specific steps of S12 are as follows: S121: Using the Lorentz model, let K = 1, then Expressed as The origin of express S122: For ZR i , mapping it into the hyperbolic space, the formula is as follows: in, Represents the drug node dr i The initial eigenvector of the hyperbolic drug, exp o (·) is used to map variables to hyperbolic space (0, ZR i )satisfy represents the Minkowski inner product; ZR i Represents the drug node dr i The initial feature vector of the drug; ||.||2 represents the L2 norm; S123: By any drug node dr i The initial drug feature vector ZR i By performing the above operations, we can obtain the initial characteristic matrix of hyperbolic drugs S124: Repeat the above steps to obtain the initial feature matrix of the hyperbolic disease The specific steps of S13 are as follows: Obtain drug-disease association data and construct the adjacency matrix A. The specific construction method is as follows:
4. The method for predicting drug-disease association based on a hyperbolic graph feature learning network according to claim 1, characterized in that: The specific steps of S2 are as follows: S21: Use hyperbolic feature transformation to extract intermediate feature vectors of drug nodes and disease nodes; S22: Calculate the hyperbolic homogeneous fusion attention and use the hyperbolic homogeneous fusion attention to extract the hyperbolic similarity features of drug and disease nodes.
5. The method for predicting drug-disease association based on a hyperbolic graph feature learning network according to claim 4, characterized in that: The specific steps of S21 are as follows: S211: For nr drug nodes in the drug-disease association data, the hyperbolic drug initial feature matrix ZR is known. H ; Indicates ZR H The initial eigenvector of the hyperbolic drug in the i-th row; At level 0, the hyperbolic drug similarity feature matrix is as follows: SR 0,H =ZR H but Indicates SR 0,H The hyperbolic drug similarity feature vector of the i-th row; S212: For any drug node dr in the drug-disease association data i , in In the layer hyperbolic feature reconstructor, firstly, the hyperbolic feature transformation is performed to map the hyperbolic space of the previous layer to the hyperbolic space of this layer. The hyperbolic feature transformation formula of the layer is as follows: in, Represented as drug node dr i After the hyperbolic characteristic transformation, the The intermediate eigenvector of ; It is The trainable parameters of the layer, It is Dimensions of layers; It is the drug node dr i In the Hyperbolic drug similarity eigenvectors of the layers; It is The layer is located at The Euclidean vector of is a trainable parameter; Represents the drug node dr i On the next floor is located in hyperbolic space Intermediate variables in exp o The role of (·) is to map the variables to the next layer Hyperbolic space log o The function of (·) is to map the variable to the tangent space of the origin o The effect is point As the center, map the variables to the hyperbolic space middle; Will The tangent space from the origin o Parallel transport to point The tangent space The specific steps of S22 are as follows: S221, calculating hyperbolic homogeneous fusion attention, where the hyperbolic homogeneous fusion attention consists of node feature attention, hyperbolic distance attention, and structural attention; S2211: Calculate the Drug node dr in the layer i For drug node dr j Node feature attention The specific formula is: Among them, LeakReLU is the activation function; is a trainable parameter; log o The function of (·) is to map the variable to the tangent space of the origin o express and Perform matrix multiplication. Then normalize to get the final node feature attention The formula is as follows: in, Represents the drug node dr i The first-order neighbor set of , including itself; when the drug node dr i and drug node dr j In the The higher the similarity of the node features of the layer, The bigger; S2212: Calculate the Drug node dr in the layer i For drug node dr j Hyperbolic distance attention The specific formula is: Among them, δ and η are hyperparameters used to adjust the hyperbolic distance; Used for calculation and The hyperbolic distance between represents the square of the hyperbolic distance; Then normalize to get the final hyperbolic distance attention The specific formula is: in, Represents the drug node dr i The first-order neighbor set of , including itself; when the drug node dr i and drug node dr j In the The smaller the hyperbolic distance of the layer, The bigger; S2213: Calculate drug node dr i For drug node dr j The structural attention of , the specific calculation process is as follows: Known drug similarity matrix DR, drug node dr i The degree of is defined here as: The formula means to sum each column j of the i-th row of DR, that is, the drug node dr i For each drug node dr j Sum the similarities; Drug Node dr i For drug node dr j Structural Attention The calculation formula is as follows: in, Then normalize to get the final structural attention The formula is as follows: in, Represents the drug node dr i The first-order neighbor set of , including itself; when the drug node dr i and drug node dr j The smaller the degree of The bigger; S2214: layer, paying attention to node features Hyperbolic distance attention and structural attention The fusion obtains the hyperbolic homogeneous fusion attention, the formula is as follows: where w NF 、w DS and w ST It is a trainable parameter and is initialized to 1; S222: Extracting hyperbolic similarity features of drug and disease nodes using hyperbolic homogeneous fusion attention; S2221: Yes Perform hyperbolic encoding to obtain the drug node dr i In the Hyperbolic similarity characteristics of layers The formula for hyperbolic encoding is as follows: in, Its function is to map variables to hyperbolic space In the middle; σ is the activation function; Represents the drug node dr i The first-order neighbor set of , including itself; The function is to map variables to The tangent space middle; S2222: For any drug node dr i Perform the above operations and iterate through h layers to finally obtain the hyperbolic drug similarity feature matrix use Indicates SR h,H ; S2223: Repeat the above steps to obtain the hyperbolic disease similarity feature matrix 6. The method for predicting drug-disease association based on a hyperbolic graph feature learning network according to claim 1, characterized in that: The specific steps of S3 are as follows: S31: Use GCN to obtain hyperbolic disease heterogeneous initial feature matrix and hyperbolic drug heterogeneous initial feature matrix; S32: Calculate the hyperbolic heterogeneous changing attention and use the hyperbolic heterogeneous changing attention to learn the heterogeneous features of drug nodes and disease nodes.
7. The method for predicting drug-disease association based on a hyperbolic graph feature learning network according to claim 6, characterized in that: The specific steps of S31 are as follows: S311: Construct a heterogeneous graph AG based on the adjacency matrix A as follows: in, S312: In order to reflect the number of layers, let MR 0,H Expressed as ZR H , MD 0,H Indicated as ZD H ; In order to use the GCN formula, MR 0,H ,MD 0,H After splicing, we get: in, E 0 It is the 0th layer by MR 0,H ,MD 0,H The concatenated feature matrix; S313: layer, obtained using GCN based on heterogeneous graph AG The formula is as follows: in, yes In the The hyperbolic heterogeneous initial eigenvector of the i-th row of the layer; σ is the activation function; exp o (·) is used to map variables to hyperbolic space It is the drug node dr i or disease node ds i The set of first-order neighbors of , including itself; is the processed adjacency matrix, yes The value of row i and column j; I is the identity matrix; is a trainable parameter; log o (·) is used to map the variable to the tangent space of the origin o The D matrix is used to calculate S314: After h layers of iteration, we can get: That is, from E h The hyperbolic drug heterogeneous initial characteristic matrix is obtained from and hyperbolic disease heterogeneous initial feature matrix The specific steps of S32 are as follows: Given the hyperbolic drug heterogeneous initial feature matrix and the hyperbolic disease heterogeneous initial feature matrix, namely MR h,H and MD h,H ; At level 0, the hyperbolic drug heterogeneous feature matrix and the hyperbolic disease heterogeneous feature matrix are: HR 0,H =MR h,H HD 0,H =MD h,H The hyperbolic heterogeneous change graph converter has three inputs: the query matrix Q 1 and Q 2 , key matrix K 1 and K 2 And the value matrix V; if the hyperbolic drug heterogeneous features are extracted, the query matrix is obtained by HR 0,H The key matrix and value matrix are obtained through HD 0,H If we extract the hyperbolic disease heterogeneous features, the situation is just the opposite; in the The outputs are and In the layer, we get The process is as follows: in, It is Hyperbolic drug heterogeneous feature matrix of the layer; exp o (·) is used to map variables to hyperbolic space log o (·) is used to map the variable to the tangent space of the origin o express and Perform matrix multiplication, Similarly; in, It is Hyperbolic disease heterogeneous feature matrix of the layer; The meaning is the same as above; Definition The drug node dr of the layer i The update equation is as follows: in, Indicates in The drug node dr of the layer i Hyperbolic drug heterogeneous knowledge feature vector; V j is the jth row of V; represents the Lorentz norm; yes Hyperbolic drug heterogeneity feature vector of row i; Norm is LayerNorm or BatchNorm; exp o (·) is used to map variables to hyperbolic space log o (·) is used to map the variable to the tangent space of the origin o RDIF ij Defined as: RDIF ij is the hyperbolic attention score, Defined as: in, It is The scalar of the layer, It is The learnable parameters of the layer, λ init is a hyperparameter; After h layers of iteration, the drug node dr i The output hyperbolic drug heterogeneity feature vector is Then we get the hyperbolic drug heterogeneous feature matrix Use HR H HR h,H ; Similarly, the hyperbolic disease heterogeneous feature matrix HD can also be obtained at the hth layer H .
8. The method for predicting drug-disease association based on a hyperbolic graph feature learning network according to claim 1, characterized in that: The specific steps of S4 are as follows: Calculate the hyperbolic drug and disease feature matrix as follows: FR H =exp o (MLP(log o (SR H )||log o (HR H ))) FD H =exp o (MLP(log o (SD H )||log o (HD H ))) Among them, exp o (·) is used to map variables to hyperbolic space log o (·) is used to map the variable to the tangent space of the origin o The role of MLP is to transform the feature vectors of hyperbolic drugs and diseases from Dimensionality reduction to 9. The method for predicting drug-disease association based on a hyperbolic graph feature learning network according to claim 1, characterized in that: The specific steps of S5 are as follows: S51: Construct drug negative samples and disease negative samples; S52: Use the positive-negative fusion hard sample sampling strategy to generate hard drug negative samples and hard disease negative samples and obtain the hyperbolic hard drug feature matrix and the hyperbolic hard disease feature matrix; The specific steps of S51 are as follows: 511: First, the hyperbolic drug feature matrix and the hyperbolic disease signature matrix Mapped into drug feature matrix XR and disease feature matrix XD, as follows: XR=log o (FR H ) XD=log o (FD H ) in, log o The function is to map the variable to the tangent space of the origin o middle; S512: Obtain drug-disease association data and generate a heterogeneous graph G, which includes drug-disease relationships, drug-drug relationships, and disease-disease relationships; find all positive sample pairs and negative sample pairs from the heterogeneous graph G. First, ensure that the number of positive sample pairs and negative sample pairs is the same. The strategy adopted is: randomly select a part of negative sample pairs from the heterogeneous graph G to ensure that the number is the same as that of positive sample pairs, and vice versa; finally, select drug negative samples and disease negative samples from the selected negative sample pairs; where the drug negative sample feature matrix is The disease negative sample feature matrix is N neg is the number of negative sample pairs, that is, the number of drug negative samples and disease negative samples; The specific steps of S52 are as follows: S521: Construct the feature matrix of hard negative candidate samples; For any node in the heterogeneous graph G, the set consisting of all first-order and second-order neighbor nodes is called the neighbor set of the node; For all drug negative samples, find all its first-order and second-order neighbors, that is, the neighbor set; For any drug negative sample r, the operation of constructing the feature matrix of hard candidate negative samples is as follows: First, find all drug positive samples in its neighbor set according to the heterogeneous graph G. The local neighbor drug positive sample feature matrix composed of its drug positive samples is: XR pl , then fuse the feature vectors of all drug positive samples to obtain the drug fusion positive sample feature vector x pr , the formula is as follows: in It's XR pl The drug positive sample feature vector of the mr-th row; mr represents the number of drug positive samples in the neighbor set of drug negative sample r; Similarly, the feature vectors of all disease positive samples in its neighbor set are fused to obtain the disease fusion positive sample feature vector x pd , the formula is as follows: in Yes XD pl The disease positive sample feature vector of the md-th row; md represents the number of disease positive samples in the neighbor set of the drug negative sample r; Then find the disease negative samples of the second-order neighbors in the drug negative sample r neighbor set, and the second-order neighbor disease negative sample feature matrix composed of them is: XD nl ; According to x pr and x pd The probability distribution of disease negative samples is calculated for the disease negative samples of the second-order neighbors in the drug negative sample r neighbor set. The formula is as follows: in, XD nl The qth row feature vector of ; μ is a hyperparameter used to balance the effects of drugs and diseases; N represents the number of disease negative samples of the second-order neighbors in the neighbor set of drug negative samples r; The meaning of this formula is the distance (x pr ,x pd )The closer the disease negative sample is, the greater the probability of being selected as a candidate disease negative sample; Then according to the probability distribution of disease negative samples Select the disease negative samples of the second-order neighbors in the neighbor set of drug negative samples r corresponding to the M largest probability values, and construct a candidate disease negative sample feature matrix of size M M is a hyperparameter; Next, the disease fusion positive sample feature vector x pd The information is injected into the candidate disease negative sample feature matrix T to construct the difficult candidate disease negative sample feature matrix T'. The formula is as follows: T' i =αx pd +(1-α)T i ,α∈(0,1) Among them, T' i is the negative sample feature vector of the hard disease in the i-th row of T'; α is a hyperparameter; T i is the negative sample feature vector of candidate disease in row i of T; S522: Obtaining a hyperbolic difficult-case drug feature matrix and a hyperbolic difficult-case disease feature matrix; For each row of the difficult candidate disease negative sample feature matrix T', the difficult disease negative sample feature vector T' j The drug negative sample feature vector of the drug negative sample r Dot product, the formula is as follows: in, It is the feature vector of the difficult disease negative sample with the largest dot product selected from T'; Since there are N negative samples of drugs neg So we can generate N neg The hard-to-detect disease negative sample feature vectors constitute the hard-to-detect disease negative sample feature matrix Similarly, for disease negative samples, N neg difficult drug negative sample feature vectors, forming a difficult drug negative sample feature matrix Finally based on XR' n and XD' n Update XR and XD to obtain the difficult drug feature matrix XR' and the difficult disease feature matrix XD', where XR' is the same size as XR, and XD' is the same size as XD; Then map it back to the hyperbolic space to get the hyperbolic difficult drug feature matrix FR' H and the hyperbolic hard-to-treat disease feature matrix FD' H ,as follows: FR' H =exp o (XR') FD' H =exp o (XD') where exp o Its function is to map variables to hyperbolic space middle.
10. The method for predicting drug-disease association based on a hyperbolic graph feature learning network according to claim 1, characterized in that: The specific steps of S6 are as follows: S61: Calculate FR' H and FD' H Hyperbolic distance ij ; Compute hyperbolic hard-to-treat drug feature vectors and hyperbolic hard-to-treat disease feature vectors Hyperbolic distance ij , the calculation formula is as follows: in, Distance ij The smaller it is, the higher the possibility of drug-disease association; S62: Using Distance, determine the predicted drug-disease association probability output through MLP; output = MLP(Distance) in,
Citation Information
Patent Citations
Drug relocation method based on multi-view stacked jump graph convolutional network
CN119905174A
Method and System for Assessing Drug Efficacy Using Multiple Graph Kernel Fusion
US20210134418A1
Cited By
Drug relocation method based on two-channel graph comparative learning
CN121964192A