Drug-disease association prediction method and system based on heterogeneous node sequence representation

Modeling biological interaction networks through heterogeneous graph neural network and sequence representation method solves the problem that existing methods are difficult to capture heterogeneous attributes and distinguish information from different sources, and achieves more accurate drug-disease association prediction.

CN119274691BActive Publication Date: 2025-05-23DALIAN MARITIME UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411344030.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-25
Publication Date
2025-05-23
Estimated Expiration
2044-09-25

AI Technical Summary

Technical Problem

Existing drug-disease association prediction methods are difficult to fully capture the heterogeneous properties of biological interaction networks, and it is difficult to distinguish neighborhood information transmitted by different sources, resulting in semantic information confusion.

Method used

The biological interaction network is modeled using heterogeneous graph neural network, transforming each node into a sequence representation, and neighborhood information obtained from different metarelations and levels is retained in different vectors in the sequence, and cross-domain information fusion is carried out through the multi-head attention mechanism and the multi-head self-attention mechanism.

Benefits of technology

It effectively avoids confusion of semantic information, learns more distinguished drug and disease vector representations, and improves the accuracy of drug-disease association prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119274691B_ABST
    Figure CN119274691B_ABST
Patent Text Reader

Abstract

The present invention provides a drug-disease association prediction method and system based on heterogeneous node sequence representation, which belongs to the field of drug discovery technology in bioinformatics. Construct a drug similarity network and a disease similarity network; construct a heterogeneous biological interaction network using the drug similarity network and the disease similarity network to obtain a first-order neighbor subgraph of a target node; split the first-order neighbor subgraph of the target node into multiple meta-relationship bipartite graphs according to the meta-relationship type, and perform intra-domain message propagation and aggregation in each meta-relationship bipartite graph to update the semantic features of the target node; fuse the semantic features of the target node across domains, and update the sequence features and aggregation features of the target node; use the updated sequence features to obtain a vector representation of the target node through a multi-head self-attention mechanism; input the vector representations of the drug node and the disease node into a multi-layer perceptron to predict the probability of drug-disease association. The present invention effectively avoids confusion of semantic information, thereby improving the accuracy of model prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of drug discovery in bioinformatics, and in particular to a drug-disease association prediction method and system based on heterogeneous node sequence representation. Background Art

[0002] The present invention focuses on the task of drug-disease association prediction, that is, predicting the probability that a drug can treat a given disease. The high-confidence drug-disease pairs output by the prediction model are recommended to researchers as candidates to achieve effective and efficient drug repositioning, thereby accelerating the drug development process.

[0003] Since the relationship between drugs and diseases can be naturally modeled as a network structure, the existing technology uses graph embedding methods such as graph representation learning and graph neural networks to project drug and disease nodes into a low-dimensional dense vector space, and then predict the potential drug-disease relationship through a neural network classifier. However, the existing methods have the following two defects in improving the prediction accuracy of the model: (1) Previous work simply assumes that different types of nodes / edges (nodes include drugs, diseases, proteins, etc., and edges include drug-disease, disease-protein, etc.) share exactly the same representation space, or separately models drug and disease similarity networks and drug-disease association networks, and simply splices the vector representations obtained from the above two networks, making them insufficient to fully capture the heterogeneous properties of the network. (2) Previous work uses multi-layer graph neural networks to learn the low-order and high-order neighborhood information corresponding to the target node. The unique vector representation corresponding to each node mixes neighborhood information from different relationships and different levels, making it difficult to further distinguish messages transmitted to the target node from different sources, resulting in semantic information confusion.

[0004] Therefore, a drug-disease association prediction method that fully captures the heterogeneous properties of the network and has clear node information is needed. Summary of the invention

[0005] In view of this, the present invention provides a drug-disease association prediction method and system based on heterogeneous node sequence representation, which uses a heterogeneous graph neural network to model the biological interaction network and converts each node into a sequence representation, thereby retaining the neighborhood information obtained from different meta-relationships and different levels, and effectively avoiding semantic information confusion.

[0006] To this end, the present invention provides the following technical solutions:

[0007] Drug-disease association prediction methods based on heterogeneous node sequence representation include:

[0008] Construct drug similarity network and disease similarity network based on similarity measurement method;

[0009] Construct heterogeneous biological interaction networks using drug similarity networks and disease similarity networks;

[0010] Based on the heterogeneous biological interaction network, the first-order neighbor subgraph of the target node is obtained;

[0011] Split the first-order neighbor subgraph of the target node into multiple meta-relation bipartite graphs according to the meta-relation type;

[0012] Perform intra-domain message propagation and aggregation in the bipartite graph of each element relationship of the target node, and update the semantic features of the target node;

[0013] The multi-head attention mechanism is used to fuse the semantic features of the target node across domains and update the sequence features and aggregation features of the target node;

[0014] Using the updated sequence features, the vector representation of the target node is obtained through a multi-head self-attention mechanism;

[0015] The target node vector representation includes a drug node vector representation and a disease node vector representation;

[0016] The vector representations of drug nodes and disease nodes are input into a multilayer perceptron to predict the drug-disease association probability.

[0017] Furthermore, the construction of the drug similarity network and the disease similarity network based on the similarity measurement method includes:

[0018] According to the drug fingerprint similarity and Gaussian interaction spectrum kernel similarity, the similarity between any drugs is calculated to select the top-K nearest neighbor drug nodes and construct the drug similarity network ε DRS ;

[0019] The similarity between any diseases is calculated based on the disease phenotype similarity and Gaussian interaction spectral kernel similarity, the top-K nearest neighbor disease nodes are selected, and the disease similarity relationship ε is constructed. DIS .

[0020] Furthermore, the construction of a heterogeneous biological interaction network using a drug similarity network and a disease similarity network includes:

[0021] Heterogeneous biological interaction networks

[0022] in, and The entity sets representing drug DR, disease DI and protein PR respectively;

[0023] ε={ε DRS , ε DIS , ε DDA , ε DRP , εDIP}, where ε DDA , ε DRP , ε DIP Represent the known drug-disease associations, drug-protein interactions, and disease-protein interactions, respectively; the interaction relationship between any two types of entities ε DDA , ε DRP , ε DIP Get from the dataset;

[0024] Define node and edge type mapping functions Assign a corresponding type to each node or edge; That is, three types of entities and five types of edges.

[0025] Furthermore, the meta-relation bipartite graph is represented as:

[0026]

[0027] Among them, R(φ(t)) is the total meta-relationships contained in the target node t type. The meta-relation corresponding to the edge e from the source node s to the target node t is defined as

[0028] Furthermore, the intra-domain message propagation and aggregation are performed in the bipartite graph of each element relationship of the target node to update the semantic features of the target node, including:

[0029] In the bipartite graph In layer l, the message propagation vector from source node s to target node t through edge e is:

[0030]

[0031] in, and represent the edge and node type specific parameter matrices respectively;

[0032] By average pooling, we aggregate information from all first-order neighbors and obtain the semantic vector corresponding to the target node t at the lth layer in the meta-relation bipartite graph:

[0033]

[0034] in, Represents the target node t in the bipartite graph of element relations The first-order neighbor nodes on .

[0035] Furthermore, the method of cross-domain fusing the semantic features of the target node using the multi-head attention mechanism to update the sequence features and aggregation features of the target node includes:

[0036] Compute semantic vectors by scaling dot product attention and The degree of correlation between

[0037] Through the multi-head attention mechanism, the value vector is weighted averaged and the residual connection term is added to calculate the updated semantic vector

[0038] The set of semantic vectors of the target node t on the lth layer of all meta-relation bipartite graphs is expressed as

[0039] Will Sequence features corresponding to the (l-1)th layer of target node t Perform splicing and update the l-th layer sequence features of the target node;

[0040] right Perform an average pooling operation and combine it with the aggregated features of the target node t at the (l-1) layer Perform residual connection and update the aggregation features of the lth layer

[0041] Furthermore, the method of using the updated sequence features to obtain the vector representation of the target node through a multi-head self-attention mechanism includes:

[0042] After L layers of message propagation and aggregation, the last layer sequence of the target node t is represented Passed into the multi-head self-attention model to generate the final vector representation of the target node

[0043] For collections All vectors in Perform an average pooling operation on it as the query vector;

[0044] Evaluate the importance of each vector in the sequence representation set by performing a scaled dot product operation between the key vector and the query vector;

[0045] Perform weighted averaging on the value vector to generate the vector representation corresponding to the target node t: Drug node vector representation Or disease node vector representation

[0046] Furthermore, the step of inputting the vector representation of the drug node and the disease node into a multi-layer perceptron to predict the drug-disease association probability includes:

[0047] Two additional combinatorial operators, vector subtraction and vector term-by-term multiplication, are introduced to integrate drug vector representation and disease vector representation, and the probability of drug u being associated with disease v is calculated through a multi-layer perceptron:

[0048]

[0049] Furthermore, the cross entropy loss function is used to optimize the drug-disease association probability prediction:

[0050]

[0051] Where S is a set of positive and negative training samples, is the true label.

[0052] The drug-disease association prediction system based on heterogeneous node sequence representation includes:

[0053] The network building module builds drug and disease similarity networks based on similarity measurement methods; integrates drug and disease similarity networks, drug-disease association networks, drug-protein and disease-protein interaction networks to build heterogeneous biological interaction networks;

[0054] The meta-relation bipartite graph splitting module obtains the first-order neighbor subgraph of the target node based on the heterogeneous biological interaction network; splits the first-order neighbor subgraph of the target node into multiple meta-relation bipartite graphs according to the meta-relation type;

[0055] The intra-domain message aggregation module propagates and aggregates intra-domain messages in the bipartite graph of each element relationship of the target node, and updates the semantic features of the target node;

[0056] The cross-domain message aggregation module uses the multi-head attention mechanism to fuse the semantic features of the target node across domains and update the sequence features and aggregation features of the target node;

[0057] The drug-disease association probability prediction module uses the updated sequence features to obtain the vector representation of the target node through a multi-head self-attention mechanism; the target node vector representation includes a drug node vector representation and a disease node vector representation; the vector representations of the drug node and the disease node are input into a multi-layer perceptron to predict the drug-disease association probability.

[0058] Advantages and positive effects of the present invention:

[0059] The present invention does not need to model the similarity network and the association network separately and simply concatenate the drug and disease vector representations obtained from the two networks, but proposes a general framework to jointly model the interaction relationship between multiple categories of biological entities contained in the heterogeneous biological interaction network, so as to better capture heterogeneous attributes. In order to avoid semantic information confusion caused by the unique vector representation corresponding to each entity, the present invention uses a sequence vector form to represent each node, so that each embedded vector in the sequence represents the neighborhood information aggregated from a specific meta-relationship at a certain level, thereby avoiding semantic information confusion and learning more discriminative drug and disease vector representations. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0061] Figure 1 Schematic diagram of a flow chart of an embodiment of the present invention.

[0062] Figure 2 Schematic diagram of a heterogeneous biological interaction network in the implementation of the present invention.

[0063] Figure 3 This is a diagram of the overall architecture of the model in an embodiment of the present invention. DETAILED DESCRIPTION

[0064] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0065] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0066] The present invention provides a drug-disease association prediction method and system based on heterogeneous node sequence representation, which constructs a heterogeneous biological interaction network based on drug-disease similarity networks and drug-disease association networks; secondly, according to the meta-relationship type, the first-order neighbor subgraph corresponding to the target node is split into multiple meta-relationship bipartite graphs, and two feature vectors, namely, sequence features and aggregation features, are defined for any node; next, in-domain message propagation and aggregation are performed in each meta-relationship bipartite graph to update the semantic vector of the target node; then, cross-domain message fusion is performed on the semantic vectors obtained from each meta-relationship bipartite graph, and the sequence features and aggregation features corresponding to the target node are updated; finally, given a drug-disease pair to be predicted, the treatment probability is output using the sequence representation corresponding to the drug and disease nodes.

[0067] Processing flow such as Figure 1 As shown, each step is further described below.

[0068] S1. Constructing heterogeneous biological interaction networks include:

[0069] definition in Represent the entity sets of drug (abbreviated as DR), disease (abbreviated as DI), and protein (abbreviated as PR); ε = {ε DRS , ε DIS , ε DDA , ε DRP , ε DIP}, where ε DRS , ε DIS Respectively represent drug and disease similarity networks, ε DDA , ε DRP , ε DIP They represent known drug-disease associations, drug-protein interactions, and disease-protein interactions, respectively.

[0070] Define node and edge type mapping functions Assign a corresponding type to each node or edge; That is, three types of entities and five types of edges.

[0071] Will The meta-relation corresponding to the edge e from the source node s to the target node t is defined as Used to convert nodes and edges into corresponding types.

[0072] Drug similarity relationship ε DRSThe construction process of is as follows: calculate the similarity between any drugs based on the drug fingerprint similarity and Gaussian interaction spectrum kernel similarity. For each drug, select the top-K nearest neighbor drug nodes to construct ε DRS Similarly, the similarity between any diseases is calculated based on the disease phenotype similarity and Gaussian interaction spectral kernel similarity, and the disease similarity relationship ε is constructed based on the top-K nearest neighbor disease nodes. DIS The interaction relationship between any two types of entities ε DDA , ε DRP , ε DIP All can be obtained from the dataset.

[0073] S2, based on the heterogeneous biological interaction network constructed in S1, given any target node t, from Get the first-order neighbor subgraph corresponding to the node in, Indicates that node t is For the above first-order neighbor subgraph, split it into multiple meta-relation bipartite graphs according to the meta-relation type. Where R(φ(t)) is the total number of meta-relations contained in the node type corresponding to t.

[0074] In this embodiment, any node t corresponds to two feature vectors, namely, sequence feature And the aggregation features The sequence features are used to store the information aggregated from each meta-relation bipartite graph in each graph neural network layer; while the aggregation features are used to propagate information to adjacent nodes during the message propagation process.

[0075] S3, for the target node t, based on the meta-relation bipartite graph split out by S2, in each meta-relation bipartite graph In this embodiment, at the 0th moment, it is assumed that the initial feature vector corresponding to the node t is Where d represents the vector dimension, and the sequence features and aggregation features are initialized as In any bipartite graph In the lth layer, the message propagation vector calculation formula from the source node s to the target node t through the edge e is:

[0076]

[0077] in, as well as They are the parameter matrices specific to edge and node types respectively.

[0078] By averaging the pooling information from all first-order neighbors, we get The semantic vector corresponding to the target node t in the upper layer l:

[0079]

[0080] in, Indicates that node t is The first-order neighbor nodes on Also a node type specific parameter matrix.

[0081] S4) When executing S3, all the element relations corresponding to the target node t are bipartite graphs Get each layer of semantic vector Afterwards, the multi-head attention mechanism is used to fully mine the interactive information between the above semantic vector sets, realize cross-domain message fusion, and update the sequence features and aggregation features corresponding to the target nodes.

[0082] In this embodiment, for the lth layer semantic vector Project them onto the query vector Key Vector Value vector Where h represents the hth attention head, and the total number of attention heads is and These are all trainable parameters.

[0083] Compute two semantic vectors using scaled dot product attention and The correlation degree between them, the scaled dot product attention calculation formula corresponding to the h-th attention head is:

[0084]

[0085] New semantic vector obtained through multi-head attention mechanism It is calculated by taking a weighted average of the value vector and adding a residual connection term:

[0086]

[0087] Where σ(.) represents nonlinear activation, β is a trainable parameter, and ‖‖ represents a concatenation operation.

[0088] Put the target node t in all the meta-relation bipartite graphs The set of new semantic vectors in the layer l generated by the multi-head attention mechanism is denoted as

[0089] Will Sequence features corresponding to the (l-1)th layer of node t Splice and get the l-th layer sequence feature of the node:

[0090]

[0091]

[0092] right Perform an average pooling operation and aggregate the features corresponding to node t at the (l-1) layer Perform residual connection to realize the aggregation feature of the lth layer Updates.

[0093] S5, after L-layer message propagation and aggregation, the last layer sequence of the target node t calculated by S4 is represented as Passed into the multi-head self-attention model to generate the final vector representation of the node

[0094] Specifically: for the collection All vectors in Perform an average pooling operation on it as the query vector:

[0095]

[0096] Here, h also represents the hth attention head, and the total number of attention heads is set to H. Mapped into key vectors and value vectors, respectively For the h-th attention head, the importance of each vector in the sequence representation set is evaluated by performing a scaled dot product operation between the key vector and the query vector:

[0097]

[0098] The value vector is weighted averaged to generate the final vector representation h corresponding to the target node t t :

[0099]

[0100] In order to effectively distinguish drug and disease type entities, the final drug and disease vector representations are recorded as When a drug-disease pair (u, v) is given, its corresponding final vector representation is Introducing vector subtraction (with the symbol ) and two additional combinatorial operators for vector term-wise multiplication (denoted by Representation) further integrates drug and disease vector representations, and calculates the probability that drug u can treat disease v through a multi-layer perceptron:

[0101]

[0102] Drug-disease association prediction is optimized using the cross-entropy loss function:

[0103]

[0104] Where S is a set of positive and negative training samples, is the true label.

[0105] The present invention also provides a drug-disease association prediction system based on heterogeneous node sequence representation, comprising:

[0106] The network building module is used to construct drug-disease similarity networks based on similarity measurement methods; integrate drug-disease similarity networks, drug-disease association networks, drug-protein and disease-protein interaction networks to construct heterogeneous biological interaction networks;

[0107] The meta-relation bipartite graph splitting module is used to obtain the first-order neighbor subgraph of the target node based on the heterogeneous biological interaction network; the first-order neighbor subgraph of the target node is split into multiple meta-relation bipartite graphs according to the meta-relation type;

[0108] The intra-domain message aggregation module is used to perform intra-domain message propagation and aggregation in the meta-relation bipartite graph and update the semantic feature vector of the target node;

[0109] The cross-domain message aggregation module uses the multi-head attention mechanism to fuse the cross-domain messages of the target node and update the sequence features and aggregation features of the target node;

[0110] The drug-disease association probability prediction module uses the updated sequence features to obtain the vector representation of the target node through a multi-head self-attention mechanism; the vector representations of the drug and disease nodes in the target node are input into a multi-layer perceptron to predict the drug-disease association probability; and the cross-entropy loss function is used to optimize the drug-disease association probability prediction.

[0111] The purpose of the present invention is to propose a drug-disease association prediction method and system based on heterogeneous node sequence representation to address the problems faced by existing methods, such as the inability to fully capture network heterogeneity and the difficulty in effectively distinguishing neighborhood information from different sources. (1) Previous work simply assumes that different types of nodes / edges share exactly the same representation space, or models similarity networks and association networks separately, which makes it impossible to fully capture the heterogeneous properties of the network; (2) the single vector representation corresponding to each entity causes the higher-level graph neural network to be unable to effectively distinguish the neighborhood information transmitted from different sources, resulting in semantic information confusion. By constructing a heterogeneous biological interaction network and jointly processing the interaction relationships between multiple heterogeneous biological entities within a common framework, the heterogeneous properties of the biological interaction network can be more effectively captured. Each node is represented in the form of a sequence vector, and different vector representations in the sequence represent the neighborhood information obtained by the target entity from different meta-relationships and different levels, thereby avoiding semantic information confusion. As a result, the accuracy of drug-disease association prediction is improved.

[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A drug-disease association prediction method based on heterogeneous node sequence representation, characterized in that: include: Construct drug similarity network and disease similarity network based on similarity measurement method; Construct heterogeneous biological interaction networks using drug similarity networks and disease similarity networks; Based on the heterogeneous biological interaction network, the first-order neighbor subgraph of the target node is obtained; Split the first-order neighbor subgraph of the target node into multiple meta-relation bipartite graphs according to the meta-relation type; Perform intra-domain message propagation and aggregation in the bipartite graph of each element relationship of the target node, and update the semantic features of the target node; The multi-head attention mechanism is used to fuse the semantic features of the target node across domains and update the sequence features and aggregation features of the target node; Using the updated sequence features, the vector representation of the target node is obtained through a multi-head self-attention mechanism; The target node vector representation includes a drug node vector representation and a disease node vector representation; The vector representations of drug nodes and disease nodes are input into a multilayer perceptron to predict the drug-disease association probability.

2. The drug-disease association prediction method based on heterogeneous node sequence representation according to claim 1, characterized in that: The method of constructing a drug similarity network and a disease similarity network based on a similarity measurement method includes: According to the drug fingerprint similarity and Gaussian interaction spectrum kernel similarity, the similarity between any drugs is calculated to select the top-K nearest neighbor drug nodes and construct the drug similarity network ε DRS ; The similarity between any diseases is calculated based on the disease phenotype similarity and Gaussian interaction spectral kernel similarity, the top-K nearest neighbor disease nodes are selected, and the disease similarity relationship ε is constructed. DIS .

3. The drug-disease association prediction method based on heterogeneous node sequence representation according to claim 2, characterized in that: The method of constructing a heterogeneous biological interaction network using a drug similarity network and a disease similarity network includes: Heterogeneous biological interaction networks in, and The entity sets representing drug DR, disease DI and protein PR respectively; ε={ε DRS ,ε DIS ,ε DDA ,ε DRP ,ε DIP }, where ε DDA , ε DRP , ε DIP Represent the known drug-disease associations, drug-protein interactions, and disease-protein interactions, respectively; the interaction relationship between any two types of entities ε DDA , ε DRP , ε DIP Get from the dataset; Define node and edge type mapping function φ: Assign a corresponding type to each node or edge; That is, three types of entities and five types of edges.

4. The drug-disease association prediction method based on heterogeneous node sequence representation according to claim 3, characterized in that: The bipartite graph of the element relation is represented as: Among them, R(φ(t)) is the total meta-relationships contained in the target node t type. The meta-relation corresponding to the edge e from the source node s to the target node t is defined as 5. The drug-disease association prediction method based on heterogeneous node sequence representation according to claim 4, characterized in that: The intra-domain message propagation and aggregation are performed in the bipartite graph of each element relationship of the target node to update the semantic features of the target node, including: In the bipartite graph In layer l, the message propagation vector from source node s to target node t through edge e is: in, and represent the edge and node type specific parameter matrices respectively; By average pooling, we aggregate information from all first-order neighbors and obtain the semantic vector corresponding to the target node t at the lth layer in the meta-relation bipartite graph: in, Represents the target node t in the bipartite graph of element relations The first-order neighbor nodes on .

6. The drug-disease association prediction method based on heterogeneous node sequence representation according to claim 5, characterized in that: The method uses a multi-head attention mechanism to fuse the semantic features of the target node across domains and update the sequence features and aggregation features of the target node, including: Compute semantic vectors by scaling dot product attention and The degree of correlation between Through the multi-head attention mechanism, the value vector is weighted averaged and the residual connection term is added to calculate the updated semantic vector The set of semantic vectors of the target node t on the lth layer of all meta-relation bipartite graphs is expressed as Will Sequence features corresponding to the (l-1)th layer of target node t Perform splicing and update the l-th layer sequence features of the target node; right Perform an average pooling operation and combine it with the aggregated features of the target node t at the (l-1) layer Perform residual connection and update the aggregation features of the lth layer 7. The drug-disease association prediction method based on heterogeneous node sequence representation according to claim 6, characterized in that: The method uses the updated sequence features to obtain the vector representation of the target node through a multi-head self-attention mechanism, including: After L layers of message propagation and aggregation, the last layer sequence of the target node t is represented Passed into the multi-head self-attention model to generate the final vector representation of the target node For collections All vectors in Perform an average pooling operation on it as the query vector; Evaluate the importance of each vector in the sequence representation set by performing a scaled dot product operation between the key vector and the query vector; Perform weighted averaging on the value vector to generate the vector representation corresponding to the target node t: Drug node vector representation Or disease node vector representation 8. The drug-disease association prediction method based on heterogeneous node sequence representation according to claim 7, characterized in that: The step of inputting the vector representation of the drug node and the disease node into a multi-layer perceptron to predict the drug-disease association probability includes: Two additional combinatorial operators, vector subtraction and vector term-by-term multiplication, are introduced to integrate drug vector representation and disease vector representation, and the probability of drug u being associated with disease v is calculated through a multi-layer perceptron:

9. The drug-disease association prediction method based on heterogeneous node sequence representation according to claim 8, characterized in that: The method further includes: optimizing the drug-disease association probability prediction using a cross entropy loss function; the cross entropy loss function is: Where S is a set of positive and negative training samples, is the true label.

10. A drug-disease association prediction system based on heterogeneous node sequence representation, characterized in that: include: The network building module builds drug and disease similarity networks based on similarity measurement methods; integrates drug and disease similarity networks, drug-disease association networks, drug-protein and disease-protein interaction networks to build heterogeneous biological interaction networks; The meta-relation bipartite graph splitting module obtains the first-order neighbor subgraph of the target node based on the heterogeneous biological interaction network; Split the first-order neighbor subgraph of the target node into multiple meta-relation bipartite graphs according to the meta-relation type; The intra-domain message aggregation module propagates and aggregates intra-domain messages in the bipartite graph of each element relationship of the target node, and updates the semantic features of the target node; The cross-domain message aggregation module uses the multi-head attention mechanism to fuse the semantic features of the target node across domains and update the sequence features and aggregation features of the target node; The drug-disease association probability prediction module uses the updated sequence features to obtain the vector representation of the target node through a multi-head self-attention mechanism; the target node vector representation includes a drug node vector representation and a disease node vector representation; the vector representations of the drug node and the disease node are input into a multi-layer perceptron to predict the drug-disease association probability.

Citation Information

Patent Citations

  • Method for predicting drug-side effect relationship based on graph neural network

    CN112216396A

  • Disease-related circular RNA (Ribonucleic Acid) identification method based on graph attention

    CN114944192A