Path-based graph model interpretable method and device
By constructing a bioinformatics heterogeneous graph and using a relational graph convolutional network and path analysis method, the problem of inaccurate interpretation in drug synergy prediction is solved, and the efficient interpretability of drug synergy prediction and the guidance of clinical application are achieved.
Patent Information
- Application Number
- CN202510817826.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-10-03
AI Technical Summary
Existing drug synergy prediction methods lack accuracy in interpreting drug combinations and ignore the functional synergy between genes, limiting the feasibility of clinical applications.
Construct a bioinformatics heterogeneous graph with multiple node types and multiple edge types, learn node embedding features through a relational graph convolutional network, use a multi-layer perceptron for drug synergy prediction, extract subgraphs through N-hop and K-core algorithms, combine mask learning and Dijkstra algorithm to find important paths, and generate explanations for drug synergy prediction.
Improved interpretability and accuracy of drug synergy predictions provide systems-level insights that can help guide clinical treatment selection.
Smart Images

Figure CN120748612A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of neural network application technology, and specifically to a path-based graphical model interpretable method and device for interpreting drug synergy prediction. Background Art
[0002] Drug combination therapy refers to the combined use of two or more drugs to treat a disease. Due to its advantages such as good therapeutic effects, reduced toxic side effects, and overcoming drug resistance, drug combinations have good prospects in the treatment of cancer. However, screening for the best synergistic anti-cancer drug combinations is a very challenging task. Due to the surge in the number of drug combinations, the use of traditional clinical-based experiments to discover drug combinations is time-consuming and labor-intensive, and may also pose safety risks to patients. In recent years, researchers have established multiple anti-cancer drug combination data sets with important research value, resulting in a large number of computational methods for predicting synergistic drug combinations.
[0003] The rapid development of machine learning technology in recent years has promoted a large number of drug synergy prediction methods based on machine learning to determine whether there is synergy between drug combinations. However, most methods do not fully explain drug synergy, and only a few methods can provide certain explanations. The interpretability of some methods is related to the attention mechanism. some The interpretable method uses Shapley values.
[0004] The above research methods focus on using a single gene to explain drug synergy prediction, which ignores the functional synergy between genes in complex networks, resulting in inaccurate interpretation of drug synergy and limiting the feasibility of clinical application. Summary of the Invention
[0005] The present application provides a path-based graphical model interpretable method and device for explaining drug synergy prediction, which has excellent performance in drug synergy prediction and is also effective in interpreting prediction results.
[0006] In a first aspect, embodiments of the present application provide a path-based graphical model interpretable method for interpreting drug synergy prediction, the path-based graphical model interpretable method comprising: Acquire drug synergy prediction related data, and construct a multi-node type and multi-edge type bioinformatics heterogeneous graph based on the acquired drug synergy prediction related data; The node embedding features are learned through a relational graph convolutional network, and the learned drug combination and cell line node embedding features are aggregated to obtain the final embedding representation; Inputting the obtained final embedding representation into a multi-layer perceptron to output the probability of synergy between drug pairs on a specific cell line, thereby realizing the construction and training of a drug synergy prediction model; Based on the trained drug synergy prediction model, synergistic drug, drug, and cell line triplets were selected from the independent test set; For each collaborative triple, the nodes in the triple are taken as target nodes, and the N-hop and K-core algorithms are used to calculate the biological information heterogeneous graph to obtain the subgraph; Mask learning is implemented on all edges of the subgraph to filter out non-important edges, and the Dijkstra algorithm is used in the subgraph after mask learning to find the path with the highest preset number of scores as the explanation for drug synergy prediction.
[0007] In conjunction with the first aspect, in one embodiment, The drug synergy prediction related data includes drug-gene association data, cell line-gene association data, pathway-gene association data, and drug synergy data; The drug-gene association data are obtained based on FDA-approved or clinically investigated drugs; The cell line-gene association data were obtained from the Cancer Cell Line Encyclopedia; The pathway-gene association data were obtained from the therapeutic target database TTD; The drug synergy data is a drug combination dataset obtained from the DrugCombDB database. The drug combination dataset contains a variety of different drug combinations, and a virtual node is set for each drug combination to represent the drug-drug combination. The drugs in the drug combination are associated with the virtual nodes through directed edges.
[0008] In conjunction with the first aspect, in one embodiment, The biological information heterogeneous graph includes multiple types of nodes and multiple types of edges; The nodes of the biological information heterogeneous graph include drugs, genes, pathways, cell lines, and virtual nodes; The edges of the biological information heterogeneous graph include virtual nodes-drug, drug-gene, pathway-gene, and cell line-gene; When processing different types of nodes and edges in biological information heterogeneous graphs, all types of nodes and edges are mapped to the same feature space. Biological entity nodes Represented as a vector , No. Relationship Type Represented as a vector ,in, Indicates the number of biological entity node types, represents the number of edge types, represents the embedding size of the biological entity node, Represents the embedding dimension of the relation type.
[0009] In conjunction with the first aspect, in one embodiment, learning the embedded features of nodes through a relational graph convolutional network specifically includes: In the bioinformatics heterogeneous graph R-GCN is used to learn the embedding features of nodes, where Represents a collection of nodes, represents the edge set, Represents a set of node types, Represents a set of edge types; In the In the R-GCN layer, for node Intermediate embedding The update is calculated as follows:
[0010] in, represents a nonlinear activation function, Representation node In the layer neural network representation, Representation node In the layer neural network representation, Indicates The neighbor set of the node under the edge type, Used to balance the influence of neighbor nodes under different relationships, Representing relationships In the The learnable weight matrix of the layer, Representation node In the first The trainable self-loop weight matrix of the layer, Representation node In the Representation of layer neural network, nodes is a node The neighboring node.
[0011] In conjunction with the first aspect, in one embodiment, inputting the obtained final embedding representation into a multilayer perceptron to output the probability of synergy between drug pairs on a specific cell line specifically includes: For the final embedding representation , which are connected in series to form a triplet Indicates that, 、 The final embedding features used to represent the two drugs, represents the final embedded features of a cell line, 、 Used to indicate two drugs, indicates cell lines; The triple The feature representation of the input multi-layer perceptron is used to calculate the score to determine the triple Is there a synergistic effect between them? The score function is expressed as:
[0012] in, A triple representing the input The probability score of collaboration between represents the activation function, represents a multilayer perceptron, ‖ represents a concatenation operator; Binary cross entropy is used as the loss function to measure the difference between the predicted results and the actual results. Specifically:
[0013] in, Represents the difference between the predicted result and the actual result, Represents a triple The total number of Represents a triple The true collaborative binary labels between them.
[0014] In conjunction with the first aspect, in one embodiment, the method of selecting a synergistic drug, drug, and cell line triplets from an independent test set based on the trained drug synergy prediction model specifically includes: Based on the trained drug synergy prediction model, the drug, drug, and cell line triplets in the independent test set are predicted to obtain the predicted synergy probability of the triplets; A threshold for binary classification is set. When the predicted synergistic probability of a triple is greater than the threshold, the current triple is judged to be a synergistic sample. When the predicted synergistic probability of a triple is not greater than the threshold, the current triple is judged to be a non-synergistic sample, thereby selecting synergistic triplets of drugs, drugs, and cell lines from the independent test set.
[0015] In conjunction with the first aspect, in one embodiment, for each collaborative triple, a node in the triple is used as a target node, and an N-hop and K-core algorithm is used to calculate the biological information heterogeneous graph to obtain a subgraph, specifically including: The three nodes in the synergistic triplet and the virtual nodes corresponding to the drug combination are defined as seed nodes, and the graph is obtained after N-hop subgraph extraction. , the figure is the union of four subgraphs generated by four seed nodes; The K-core of a biological information heterogeneous graph is defined as the node with the minimum degree. The unique largest subgraph of the biological information heterogeneous graph is K-core, which is obtained by recursively removing the graph with a degree less than , until all nodes in the biological information heterogeneous graph have a degree of at least , in order to achieve the purpose of pruning, and set the nodes in the triplet not to participate in pruning; In the figure Use K-core pruning to obtain the final computational subgraph , thereby extracting a computational subgraph for each collaborative triple .
[0016] In combination with the first aspect, in one embodiment, performing mask learning on all edges of the subgraph to filter out non-important edges specifically includes: set up Indicates the relationship type is All edges of Is the relationship type The learnable mask of all edges of , then the relationship type is Applying a mask on the edge set of Represented as , then the edge set that implements the mask on all edges of all relation types on the subgraph is represented as:
[0017] in, represents the set of edges that enforces a mask on all edges of all relation types on the subgraph, represents the edge, represents the learned mask, represents element-wise product, Represents a set All relationship types in Perform a union operation; Using loss function and , measure the edges that have an impact on the output results of the drug synergy prediction model and the edges that can form meaningful paths; Among them, for the loss function , specifically:
[0018] in, Represents minimizing the prediction loss between the original image and the mask image, Represents the drug synergy prediction model for triples on the original graph The predictions are collaborative, represents the mask image, represents the probability of predicting a triplet as a synergistic combination on the mask map; Among them, for the loss function , specifically:
[0019] in, represents maximizing the weight of the edges on the explanation path, 、 represents the hyperparameter that balances the contributions of the two edges, represents a learnable mask for all edges of all types, represents a candidate edge, represents a non-candidate edge, Indicates type The edge, express learnable masks of type edges, Indicates type The learnable mask of the edges corresponds to the edge subset The weight value of .
[0020] In conjunction with the first aspect, in one embodiment, the method of using the Dijkstra algorithm to find a preset number of pathways with the highest scores as an explanation for drug synergy prediction specifically includes: Combined Edges The ability to influence the prediction results and the degree of the target node, calculate the edge Importance score:
[0021] in, Represents an edge The importance score of Represents an edge The degree of the target node in , Represents an edge The probability of being selected as the explanatory path, Represents an edge The mask weight of In a biological information heterogeneous graph, a path is formed by a set of edges, and the score of the path is equivalent to the sum of the scores of each edge included in the path. Specifically:
[0022] in, represents the path score, Indicates the path, Represents an edge The probability of being included in the explanatory path, ; Based on the calculated path score, combined with the shortest path algorithm Dijistra, and using the negative score of the edge As the score of each edge, a preset number of top-ranked shortest paths are obtained as target paths, thereby obtaining an explanation for drug synergy prediction.
[0023] In a second aspect, an embodiment of the present application provides a path-based graphical model interpretable device, the path-based graphical model interpretable device comprising: A data acquisition and heterogeneous graph construction module is used to acquire drug synergy prediction related data and construct a multi-node and multi-edge bioinformatics heterogeneous graph based on the acquired drug synergy prediction related data; A graph representation learning module, which is used to learn node embedding features through a relational graph convolutional network and aggregate the learned drug combination and cell line node embedding features to obtain the final embedding representation; a model training module, which is used to input the obtained final embedding representation into a multi-layer perceptron to output the probability of synergy between drug pairs on a specific cell line, thereby realizing the construction and training of a drug synergy prediction model; A synergistic triplet selection module is used to select synergistic drug, drug, and cell line triplets from an independent test set using a trained drug synergy prediction model; The subgraph extraction module is used to extract the subgraph from each collaborative triplet, taking the nodes in the triplet as the target nodes and applying the N-hop and K-core algorithms to the biological information heterogeneous graph; The mask learning and explanation path generation module is used to implement mask learning on all edges of the subgraph to filter out non-important edges, and use the Dijkstra algorithm in the subgraph after mask learning to find the path with the highest preset number of scores as the explanation for drug synergy prediction.
[0024] The beneficial effects of the technical solutions provided in the embodiments of the present application include: The drug synergy mechanism is analyzed through a path-based graph model interpretable method. Specifically, a heterogeneous graph containing multiple nodes and multiple relationships is constructed through the connections between drugs, genes, pathways, and cell lines. Then, the local structural information and relationship semantics in the graph are captured based on the relationship graph neural network to learn the embedded features of the nodes. Next, a corresponding explanation is generated for each synergistic triple. Specifically, SDCInterpreter determines a subgraph based on the instance, and then applies mask learning on the subgraph to identify important edges. Finally, the shortest path algorithm is used to find several optimal paths to analyze the drug synergy mechanism. It has excellent performance in drug synergy prediction and is also effective in the interpretability of the prediction results. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 This is a flowchart of the path-based graphical model interpretability method of this application; Figure 2 This is a schematic diagram of the functional modules of the path-based graphical model interpretable device of this application. DETAILED DESCRIPTION
[0026] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0027] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0028] First, an embodiment of the present application provides a path-based graph model interpretable method, which predicts and explains drug synergy by proposing a path-based graph model, namely SDCInterpreter. SDCInterpreter first constructs a heterogeneous graph through the associations between drugs, genes, pathways and cell line entities, and then uses a relational graph convolutional network to learn the node embedding features of the heterogeneous graph. The node embedding features are then input into a multi-layer perceptron to achieve drug synergy prediction. At the same time, in order to explore the interpretability of drug synergy prediction results, SDCInterpreter transforms the interpretation task of drug synergy prediction into the problem of finding paths connecting drug nodes and cell line nodes. Experimental results show that SDCInterpreter has excellent performance in drug synergy prediction and also shows effectiveness in the interpretability of prediction results.
[0029] First, it should be noted that existing methods for drug synergy prediction can be roughly divided into two types: phenotypic-based methods and network-based methods. Phenotypic-based methods use biochemical information and molecular properties as features to predict drug synergy. Network-based methods construct an association graph using biological entity relationships related to drug combinations and cell lines, and then use graph neural networks to obtain node embedding features from the association graph.
[0030] In GNNs (graph neural networks), interpretability refers to the ability to understand and explain the basis on which the model makes decisions or predictions. Existing research on the interpretability of graph models primarily focuses on identifying meaningful substructures or features that play an important role. While identifying important individual genes can be informative, in order to gain system-level insights into the drug synergy prediction process, the method proposed in this application differs from previous studies. This application introduces pathway information into the graph, generating pathway-based explanations for drug synergy prediction results to analyze how genes in certain pathways function. This system-level insight process helps guide clinical treatment choices. Specifically, the SDCInterpreter of this application consists of two parts: one for training a drug synergy prediction model, and the other for interpreting drug synergy predictions using a pathway-based approach.
[0031] In one embodiment, referring to Figure 1 , Figure 1 This is a flow chart of the path-based graph model interpretable method of this application. Figure 1 As shown in Figure 2, path-based graph model interpretability methods include: S1: Obtain drug synergy prediction related data, and construct a multi-node type and multi-edge type bioinformatics heterogeneous graph based on the obtained drug synergy prediction related data; That is, data preparation is carried out to collect various types of drug synergy prediction related data; then, a heterogeneous graph is constructed based on the collected data to build a bioinformatics heterogeneous graph with multiple node types and multiple relationship types; S2: Learn the node embedding features through the relational graph convolutional network, and aggregate the learned drug combination and cell line node embedding features to obtain the final embedding representation; That is, graph representation learning is performed, and the embedding features of nodes are learned using the relational graph convolutional network; S3: Inputting the obtained final embedding representation into a multi-layer perceptron to output the probability of synergy between drug pairs on a specific cell line, thereby realizing the construction and training of a drug synergy prediction model; That is, the model is trained, the learned drug combination and cell line node embedding features are aggregated to obtain the final embedding representation, and then it is input into a multi-layer perceptron to output the probability of synergy between drug pairs on a specific cell line; S4: Based on the trained drug synergy prediction model, synergistic drug, drug, and cell line triplets were selected from the independent test set; That is, synergistic triplets are selected, and based on the above-mentioned trained drug synergy prediction model, synergistic drug, drug, and cell line triplets are selected from the independent test set; S5: For each collaborative triple, the nodes in the triple are used as target nodes, and the N-hop and K-core algorithms are used to calculate the biological information heterogeneous graph to obtain the subgraph; That is, subgraph extraction is performed. For each collaborative triple, the node in the triple is used as the target node, and the N-hop and K-core algorithms are used to calculate the bioinformation heterogeneous graph to obtain the subgraph, so as to eliminate the irrelevant nodes in the bioinformation heterogeneous graph; S6: Mask learning is performed on all edges of the subgraph to filter out non-important edges, and the Dijkstra algorithm is used in the subgraph after mask learning to find the path with the highest preset number of scores as the explanation for drug synergy prediction.
[0032] That is, mask learning and explainable path generation are performed. Mask learning is implemented on all edges of the subgraph to filter out unimportant edges, and then explanation paths are generated. The Dijkstra algorithm is used on the masked biological information heterogeneous graph to find several paths with the highest scores as explanations for drug synergy prediction.
[0033] In this application, for data preparation, genes play a vital role in metabolism, signal transduction and gene regulatory networks, and are associated with many biological molecules in biological networks. Abnormal expression of genes can lead to cell line pathology. Therefore, this application collects various data related to genes in the drug synergy process, including drug-gene association data, cell line-gene association data, pathway-gene association data, and drug synergy data.
[0034] Among them, drug-gene association data are based on FDA (U.S. Food and Drug Administration) approved or clinical research drugs, including 4428 drugs, 2256 human genes, and 15051 drug-gene associations; The cell line-gene association data were obtained from the Cancer Cell Line Encyclopedia, a large-scale cell line dataset that calculates the mean, standard deviation, and z-score of all cell line gene expression data to indicate whether there is an association between cell lines and genes, namely, cell line-protein association. The cell line-protein association includes 18,022 protein-coding genes, 1,035 cancer cell lines, and 749,551 associations. The pathway-gene association data were obtained from the Therapeutic Target Database (TTD), which provides comprehensive information on the druggability characteristics of 426 successful, 1014 clinical trial, 212 preclinical / patent, and 1479 literature-reported targets, including 8531 pathway-gene associations. Among them, the drug synergy data (i.e., the customized virtual node-drug association) is a drug combination dataset obtained from the DrugCombDB database. The drug combination dataset contains a variety of different drug combinations, and a virtual node is set for each drug combination to represent the drug-drug combination. The drugs in the drug combination are associated with the virtual nodes through directed edges.
[0035] In this application, for the construction of a heterogeneous graph, in order to construct a bioinformatics heterogeneous graph that can simulate the complex interactions between drug combinations and cell line entities, this application uses the above-mentioned acquired data to construct a heterogeneous graph, which involves four types of entity association data.
[0036] The bioinformation heterogeneous graph includes multiple types of nodes and multiple types of edges. Specifically, the bioinformation heterogeneous graph consists of five types of nodes, namely drugs, genes, pathways, cell lines, and virtual nodes, and four types of edges, namely virtual nodes-drugs, drugs-genes, pathways-genes, and cell lines-genes.
[0037] When processing different types of nodes and edges in biological information heterogeneous graphs, all types of nodes and edges are mapped to the same feature space. Biological entity nodes Represented as a vector , No. Relationship Type Represented as a vector ,in, Indicates the number of biological entity node types, represents the number of edge types, represents the embedding size of the biological entity node, Represents the embedding dimension of the relation type.
[0038] In this application, for graph representation learning, considering the impact of various relationship types on node representation in biological information heterogeneous graphs, this application R-GCN is used to learn the embedded features of nodes, that is, the embedded features of nodes are learned through the relational graph convolutional network, specifically including: S201: In biological information heterogeneous graphs R-GCN (relational graph convolutional network) is used to learn the embedded features of nodes, where Represents a collection of nodes, represents the edge set, Represents a set of node types, Represents a set of edge types; R-GCN is based on the multi-relationship structure of biological information heterogeneous graphs and can explicitly model the impact of different relationships on nodes to better learn node representations; S202: In In the R-GCN layer, for node Intermediate embedding The update is calculated as follows:
[0039] in, represents a nonlinear activation function (such as the LeakyReLU activation function), Representation node In the layer neural network representation, Representation node In the layer neural network representation, Indicates The neighbor set of the node under the edge type, Used to balance the influence of neighbor nodes under different relationships, that is It is a normalization constant that balances the influence of neighbor nodes under different relationships in a specific problem, which can be learned and selected in advance. Representing relationships In the The learnable weight matrix of the layer, Representation node In the first The trainable self-loop weight matrix of the layer, Representation node In the Representation of layer neural network, nodes is a node The neighboring node.
[0040] In this application, for model training, drug synergy prediction aims to predict drug-drug-cell line triplets. A synergistic score is calculated to determine whether there is a synergistic effect between them. After graph representation learning, the final embedding features of two drugs and a cell line are obtained, which are recorded as 、 、 , then concatenate them to form a triple Indicates that the input is into the multi-layer perceptron to calculate the score to determine the triple Is there a synergistic effect between them?
[0041] The final embedding representation is then fed into a multilayer perceptron to output the probability of synergy between drug pairs on a specific cell line, specifically including: S301: For the final embedding representation obtained , which are connected in series to form a triplet Indicates that, 、 The final embedding features used to represent the two drugs, represents the final embedded features of a cell line, 、 Used to indicate two drugs, indicates cell lines; S302: triples The feature representation of the input multi-layer perceptron is used to calculate the score to determine the triple Is there a synergistic effect between them? The score function is expressed as:
[0042] in, A triple representing the input The probability score of collaboration between represents the activation function, represents a multilayer perceptron, ‖ represents a concatenation operator; S303: Binary cross entropy is used as the loss function to measure the difference between the predicted result and the actual result. Specifically:
[0043] in, Represents the difference between the predicted result and the actual result, Represents a triple The total number of Represents a triple The true collaborative binary labels between them.
[0044] In this application, the selection of synergistic triplets, i.e., based on the trained drug synergy prediction model, synergistic drug, drug, and cell line triplets are selected from the independent test set, specifically including: S401: Based on the trained drug synergy prediction model, predict the drug, drug, and cell line triplets in the independent test set to obtain the predicted synergy probability of the triplets; S402: A binary classification threshold is set. When the predicted synergy probability of a triple is greater than the threshold, the current triple is determined to be a synergistic sample. When the predicted synergy probability of a triple is less than the threshold, the current triple is determined to be a non-synergistic sample. This allows synergistic drug, drug, and cell line triplets to be selected from the independent test set. In practical applications, the binary classification threshold can be set to 0.5.
[0045] In this application, for the extraction of subgraphs, in order to capture the local domain information in the biological information heterogeneous graph and reduce the complexity of the biological information heterogeneous graph, the biological information heterogeneous graph is extracted. N-hop subgraph extraction and K-core pruning operations were performed. N-hop subgraph extraction of biological information heterogeneous graph refers to the set of all nodes that can be reached through n hops from the seed node.
[0046] That is, for each collaborative triple, the node in the triple is used as the target node, and the N-hop and K-core algorithms are used to calculate the biological information heterogeneous graph to obtain the subgraph, specifically including: S501: The three nodes in the synergistic triplet and the virtual nodes corresponding to the drug combination are defined as seed nodes, and the graph is obtained after N-hop subgraph extraction. , the figure is the union of four subgraphs generated by four seed nodes; S502: Define the K-core of the biological information heterogeneous graph as the one with the minimum node degree The unique largest subgraph of the biological information heterogeneous graph is K-core, which is obtained by recursively removing the graph with a degree less than , until all nodes in the biological information heterogeneous graph have a degree of at least , in order to achieve the purpose of pruning, and set the nodes in the triplet not to participate in pruning; S503: In the picture Use K-core pruning to obtain the final computational subgraph , thereby extracting a computational subgraph for each collaborative triple .
[0047] In this application, mask learning is to learn a mask matrix with the same dimension as the input tensor. The matrix selectively filters the original tensor through element-by-element multiplication, thereby retaining the meaningful parts. In order to discover the edges in the biological information heterogeneous graph that have an important impact on the drug synergy prediction results, this application learns a mask covering all edges of all relationship types.
[0048] That is, mask learning is implemented on all edges of the subgraph to filter out non-important edges, including: S601: Settings Indicates the relationship type is All edges of Is the relationship type The learnable mask of all edges of , then the relationship type is Applying a mask on the edge set of Represented as , then the edge set that implements the mask on all edges of all relation types on the subgraph is represented as:
[0049] in, represents the set of edges that enforces a mask on all edges of all relation types on the subgraph, represents the edge, represents the learned mask, represents element-wise product, Represents a set All relationship types in Perform a union operation; S602: Using loss function and , measure the edges that have an impact on the output results of the drug synergy prediction model and the edges that can form meaningful paths; In the explanation of drug synergy prediction, the purpose of mask learning is to set different weights for edges of different importance. Specifically, the important edges are set to high weights and the unimportant edges are set to low weights according to the learned mask. In order to make the prediction explanation generated by the learned mask more convincing, the importance of edges should be analyzed from two perspectives: the edges that affect the model output and the edges that can form meaningful paths are important. Therefore, this application adopts the loss function and Let's measure these two perspectives.
[0050] Loss Function The purpose is to learn which edges affect the model's predictive ability. If some edges in the biological information heterogeneous graph produce significant changes in the model's prediction results after being removed, then the removed edges are important and should be retained in the graph. This indicates that the graph after the masking operation should be the same as the initial Figure 1 In this way, sufficient information can be provided for the prediction of drug synergy. This is achieved by minimizing the loss between the original image and the mask image prediction results.
[0051] Among them, for the loss function , specifically:
[0052] in, Represents minimizing the prediction loss between the original image and the mask image, Represents the drug synergy prediction model for triples on the original graph The predictions are collaborative, represents the mask image, represents the probability of predicting a triplet as a synergistic combination on the mask map; Loss Function It is used to learn to select edges that can form meaningful paths. The main idea is to first determine a set of candidate edges , the paths formed by these edges are concise and contain rich information, and then through optimization To increase candidate edges The mask weights in , and reduce the non-candidate edges The mask weight of .
[0053] Among them, for the loss function , specifically:
[0054] in, represents maximizing the weight of the edges on the explanation path, 、 represents the hyperparameter that balances the contributions of the two edges, represents a learnable mask for all edges of all types, represents a candidate edge, represents a non-candidate edge, Indicates type The edge, express learnable masks of type edges, Indicates type The learnable mask of the edges corresponds to the edge subset The weight value of .
[0055] In this application, for the generation of paths, a node with a high degree in the biological information heterogeneous graph is usually a common node that is connected to many other attribute nodes. If such a node is included in the path, the path may contain insufficient information and be not concise. Therefore, a suitable path should not contain high-degree nodes. middle, It is a path that provides a lot of information and is concise.
[0056] That is, the Dijkstra algorithm is used to find the preset number of pathways with the highest scores as the explanation for drug synergy prediction, specifically including: S511: Combined Edges The ability to influence the prediction results and the degree of the target node, calculate the edge Importance score:
[0057] in, Represents an edge The importance score of Represents an edge The degree of the target node in , Represents an edge The probability of being selected as the explanatory path, Represents an edge The mask weights are obtained through learning; S512: In a biological information heterogeneous graph, a path is formed by a set of edges. The score of the path is equal to the sum of the scores of each edge included in the path. Specifically:
[0058] in, represents the path score, Indicates the path, Represents an edge The probability of being included in the explanatory path, ; S513: Based on the calculated path score, combined with the shortest path algorithm Dijistra, and using the negative score of the edge As the score of each edge, a preset number of top-ranked shortest paths are obtained as target paths, thereby obtaining an explanation for drug synergy prediction.
[0059] There are two characteristics of paths with higher scores: first, the edges in the path have a significant impact on the prediction results, and second, the node degrees in the path are relatively low. Finding the path with the highest score can be done by using the shortest path algorithm Dijistra, which uses the negative scores of the edges. As the score of each edge, the top several shortest paths found according to the above calculation method are used as the target paths to be found.
[0060] Furthermore, in order to train a high-quality explanation mask, we can combine and These two loss functions are used as the total loss function to jointly optimize mask learning.
[0061] The following is a detailed description of the path-based graph model interpretability method of this application in combination with corresponding experiments.
[0062] In the experiment, the experimental setup is first explained, and then the performance of the model SDCInterpreter in predicting drug synergy and its ability to explain drug synergy in biological networks are evaluated. Then, the effectiveness of each component in the model is explored and hyperparameter sensitivity analysis is performed. Finally, a case study is conducted to further verify the practical application value of the model.
[0063] For the experimental setup, in order to construct the bioinformatics heterogeneous graph, data were obtained from several public databases. The drug synergy dataset: drug synergy data was collected from DrugCombDB, the largest comprehensive database of drug combinations to date. DrugCombDB integrates drug synergy data from other public databases, high-throughput screening experiments, and literature. Each sample consists of two drugs and a cell line and the synergy score between them. The original data contains 69,436 drug combinations, including 764 drugs and 76 cancer cell lines. Data related to the bioinformatics heterogeneous graph: drug-gene associations were collected from Cheng's study, in which all drugs are FDA-approved. Pathway-gene associations were obtained from TTD. Cell line-gene associations were obtained from CCLE. In summary, the bioinformatics heterogeneous graph contains the following node types: 2462 virtual nodes, 571 drug 1s, 83 drug 2s, 13480 genes, 371 pathways, and 76 cell lines, and relationship types: 2462 virtual node-to-drug associations, 4436 drug 1-to-gene associations, 1283 drug 2-to-gene associations, 27730 cell line-to-gene associations, and 6713 pathway-to-gene associations.
[0064] The predictions and explanations of SDCInterpreter are compared with several state-of-the-art baseline methods to demonstrate the effectiveness of SDCInterpreter’s predictions and explanations. These baselines can be divided into the following categories: XGboost is a scalable, efficient, and classic machine learning method that performs well in various classification tasks. DeepSynergy uses a three-layer feed-forward neural network to predict synergy scores, which considers gene expression as cell line characteristics and three types of chemical descriptors as drug characteristics; GraphSynergy uses a spatial graph-based convolutional network and attention mechanism to encode high-order structural information of drug and cell line protein modules to enrich entity embedding representations; KGANSynergy uses an attention mechanism to learn node representations on heterogeneous graphs to predict drug synergy. HypergraphSynergy is a multi-directional relation-enhanced hypergraph representation learning method for predicting drug-drug cell line associations.
[0065] To evaluate the performance of SDCInterpreter, the dataset was randomly split into a cross-validation set and an independent test set in a 9:1 ratio. Then, 5-fold cross-validation was performed on the cross-validation set to optimize model hyperparameters and train the model. The independent test set was used to predict the model performance. The cross-validation set can be set in the following three modes: Random cross-validation set: To evaluate the ability of the SDCInterpreter model to rediscover known anticancer drug combinations, all samples were randomly divided into five mutually exclusive subsets of equal size to implement 5-CV; Cell line-level cross-validation set: To evaluate the predictive ability of the SDCInterpreter model for new cell lines, i.e., cell lines in the independent test set that did not appear in the training phase, 5-CV was achieved by randomly splitting samples at the cell line level; Cross-validation set at the drug combination level: In order to evaluate the predictive ability of the SDCInterpreter model for new drug combinations, that is, drug combinations in the independent test set did not appear in the training phase, the samples were randomly split at the drug combination level to achieve 5-CV.
[0066] Since drug synergy is a classification task, this application sets the threshold to 0 to binarize the synergy score to generate synergistic and non-synergistic samples. Synergistic effects greater than 0 are classified as synergistic samples, while those with negative effects are classified as non-synergistic samples. Four evaluation metrics commonly used in classification tasks were used to evaluate model performance, including accuracy (ACC), area under the curve (AUC), F1 score (F1), and area under the precision-recall curve (AUPR). To train the model, the Adam optimizer with a learning rate of 0.0001 was used to optimize the model. The best-performing model based on the area under the curve (AUC) was selected on a validation set with 1000 traversal units for testing on an independent test set. The average metric of 5-CV was used as the final result.
[0067] For performance comparison, the performance of the model and the baseline on the DrugCombDB dataset is compared. SDCInterpreter is compared with the baseline method on the DrugComDB dataset. According to the relevant results, SDCInterpreter achieves the best performance in all indicators of the CV set, proving its effectiveness in drug synergy prediction. In addition, it is observed that: (1) compared with the XGboost machine learning model that only considers entity feature information, the performance of SDCInterpreter is effectively improved, which shows that the structural information learned by the heterogeneous graph enhances the representation ability of the node; (2) compared with the deepsynergy model that only considers the feature information of drugs and cell lines, SDCInterpreter considers the complex topological structure of the heterogeneous graph, enabling it to learn the interaction information more comprehensively; (3) GraphSynergy and HypergraphSynergy methods respectively obtain topological structure information and use hypergraphs to represent triple associations. Compared with the first two methods, the performance on most metrics is significantly improved. However, SDCInterpreter not only considers the topological domain structure, but also pays attention to the relationship information between entities. This method updates the representation of nodes more comprehensively.
[0068] For ablation analysis, in order to demonstrate the importance of components in the model, several different variants of the model were designed and further compared on an independent test set.
[0069] The SDCInterpreter without pathway nodes (W / o P) removed the association between pathways and genes during the mapping process; SDCInterpreter without message passing (W / o MP) uses the initial embedding that does not contain neighbor node information; SDCInterpreter without relation weights (W / o RW) uses the same weight matrix for all relation types; SDCInterpreter without relationship type (W / o EW) only considers node embeddings but not the relationships between nodes; According to the relevant experimental results, the performance of all variants of SDCInterpreter decreased, which shows that these components are helpful for the prediction of drug synergy. Specifically, the following results can be observed: Among all variants, SDCInterpreter (W / oP) has the most significant performance degradation except for AUC, which is higher than SDCInterpreter (w / o-EW). This indicates the importance of pathway relationship information in drug synergy prediction. Pathway information can capture the potential mutual information between nodes, and removing pathway nodes will reduce the model's ability to represent nodes. Compared with SDCInterpreter, the performance of SDCInterpreter (w / o-MP) has declined, with the most significant drop in F1. This indicates that using message passing to aggregate information from neighboring nodes helps the model distinguish between positive and negative samples more accurately when dealing with datasets with unbalanced drug combinations. SDCInterpreter (w / o-RW) shows performance degradation on all evaluation metrics, which indicates that differential treatment of different types of biological relationships contributes differently to model predictions. Differentiating different types of biological relationships is beneficial to comprehensively learning the representation of nodes in the graph. All indicators of SDCInterpreter (w / o-EW) also decreased, with ACC changing the most. This suggests that jointly learning node representations using the embedding information of gene-node relationships provides useful domain knowledge for drug synergy prediction.
[0070] For case analysis, SDCInterpreter automatically assigns contribution weights to each edge by applying mask learning on the edges of the heterogeneous graph. The weight of each edge represents its contribution to synergistic prediction and its ability to contain information. Based on the edge weight, several paths to the cell line can be found for each drug in the synergistic combination. The important explanatory paths are defined as the top five paths found for each drug.
[0071] The following will explain the explainability of these important pathways from a pharmacological perspective.
[0072] In an independent test set, we selected a clinically validated synergistic drug combination, Vemurafenib and Celecoxib, in the MDAMB231 cell line. Recent studies have confirmed the therapeutic efficacy of both drugs against breast cancer. For this example, SDCInterpreter predicted a positive relationship between them. The top five pathways generated by SDCInterpreter for Vemurafenib and Celecoxib are shown in Table 1 below. Each pathway includes information on the drug and cell line, along with the drug's target, pathway, and cell line's target. The results indicate that Vemurafenib targets two proteins, RAF1 and BRAF. Both RAF1 and BRAF belong to the RAF protein family, which plays a key role in protein kinase pathways. Overactivation of RAF proteins, particularly RAF1 and BRAF, drives tumor progression and drug resistance in many cancer types. The drug Celecoxib has three targets, namely PTGS2, MAPK14, and PDPK1 genes. Among them, the MAPK14 gene has the function of regulating multiple cells. Its upregulation will lead to accelerated proliferation and migration of MDAMB breast cancer cells
[31] . PDPK1 can enhance the stability of HIF-1α protein by phosphorylating hypoxia-inducible factor 1α, and promote HIF-1α transcriptional activity by enhancing the binding of HIF-1α to P300. PDK1 can form a positive feedback loop with it, which is of great significance in promoting the development of breast cancer. In the treatment of breast cancer, the PI3K-Akt signaling pathway and Pathways in cancer play an important role. The PI3K-Akt signaling pathway can regulate cell survival, metastasis, and metabolism. Pathways in cancer integrates multiple molecular mechanisms and signal transduction pathways related to cancer pathogenesis, including the PI3K-Akt signaling pathway. Existing studies have shown that RAF is interconnected with the PI3K-Akt signaling pathway cascade feedback loop, and that PDPK1 is able to control the expression of upstream PI3K and activation of ATK, suggesting that these relationships may be potential therapeutic mechanisms. In cell lines, the target gene GSK3B (a serine) has been shown to be a key regulator of breast cancer drug resistance, participating in the signaling cascade of the PI3K-Akt signaling pathway. Loss of GSK-3β kinase activity significantly increases drug and hormone resistance in breast cancer cells.The results of SDCInterpreter showed that Vemurafenib and Celecoxib jointly regulated the PI3K-Akt signaling pathway by inhibiting RAF enzyme and PDK1 enzyme respectively to reduce the activity of the cell line target GSK3B gene, thereby generating better drug resistance and achieving synergistic treatment.
[0073] Table 1 Interpretable pathways selected by SDCInterpreter to explain drug synergy
[0074] The SDCInterpreter model was applied to the drug combination of fulvestrant and theophylline and the cell line COLO205. The top five results between the two drugs and the cell line are shown in Table 1, which summarizes the interaction between the drug combination of fulvestrant and theophylline and the cell line COLO205. Specifically, SDCInterpreter generated two explanatory pathways for each drug, encompassing the proteoglycan pathway in cancer and four target genes: ESR1, HIF1A, RAF1, and MAPK12. The COLO205 cell line is a human colon cancer cell line. Colonic malignancies are associated with a low-fiber diet. Proteoglycans, including syndecan-1, versican, and decorin, generally function by covalently forming macromolecules to inhibit tumor growth. Among these targets, ESR1 can activate decorin glycans, influencing changes in the intracellular environment, while HIF1A can activate glycan synthase under certain circumstances. The genes RAF1 and MAPK12 can play a key role in organ proliferation and regulate cell function, respectively. This shows that the two drugs affect RAF1 and MAPK12 by regulating ESR1 and HIF1A, thereby causing the degradation of lactose and achieving the therapeutic purpose.
[0075] This application proposes a path-based graphical model interpretability method to predict drug synergy and analyze drug synergy predictions from the perspective of pathway regulation. This method breaks the limitation of previous studies that focused on single-gene analysis. SDCInterpreter constructs a heterogeneous graph that links drugs, genes, pathways, and cell lines through their connections. It introduces a virtual node for each drug combination, facilitating interpretable analysis of drug synergy from a path-based perspective. To better integrate the interaction information between nodes in the heterogeneous graph, R-GCN (Relational Graph Convolutional Network) is used to learn node embedding features on the heterogeneous graph to construct a drug synergy prediction model. To explain drug synergy prediction, a path-based method is used on the heterogeneous graph to provide a certain level of interpretability. Specifically, a mask is first learned for all edge types in the heterogeneous graph through mask learning. Next, a shortest path algorithm is used to generate several important paths in the masked graph. Finally, drug synergy is explained by analyzing these path information.
[0076] In drug synergy prediction, a bioinformatics network is constructed by integrating biomedical entity data associated with drugs and cell lines. The network is a directed heterogeneous graph. ,satisfy For each node , through the mapping function Associated with a biomedical entity object type, for each edge , through the mapping function Associated with an inter-object relationship type. Node embedding features obtained using R-GCN on a heterogeneous graph are input into a multi-layer perceptron to build a drug synergy prediction model. After training, this drug synergy prediction model can determine whether drug-drug-cell line triplets have drug synergy.
[0077] A pathway-based approach is used to explain drug synergy prediction, formally representing the pathway as ,in, Indicates the node type, Indicates the type of relationship between nodes. If the starting node of the path It is a drug combination, the termination node The goal of this application is to generate several important explanatory pathways for synergy prediction between drug combinations and cell lines: Given a trained drug synergy prediction model, a subgraph determined by drug combinations and cell lines, explain the number of edges contained in the path and the number k; find the explanation path Path={p|p is the explanation path between the synergistic drug combination and the cell line, where the length of p is no greater than}, and the number of Paths is no more than k; by optimizing p∈Path, the generated paths are made concise, rich in information and have an impact on the results of drug synergy prediction.
[0078] The path-based graph model interpretable method of the embodiment of the present application analyzes the drug synergy mechanism through the path-based graph model interpretable method. Specifically, a heterogeneous graph containing multiple nodes and multiple relationships is constructed through the connections between drugs, genes, pathways, and cell lines. Then, the local structural information and relationship semantics in the graph are captured based on the relationship graph neural network to learn the embedded features of the nodes. Then, a corresponding explanation is generated for each synergistic triple. Specifically, SDCInterpreter determines a subgraph based on the instance, and then applies mask learning on the subgraph to identify important edges. Finally, the shortest path algorithm is used to find several optimal paths to analyze the drug synergy mechanism. It has excellent performance in drug synergy prediction and is also effective in the interpretability of the prediction results.
[0079] In a second aspect, an embodiment of the present application also provides a path-based graph model interpretable device.
[0080] In one embodiment, referring to Figure 2 , Figure 2 This is a functional module diagram of the path-based graph model interpretable device of this application. Figure 2 As shown, the path-based graph model interpretable device includes: a data acquisition and heterogeneous graph construction module, a graph representation learning module, a model training module, a collaborative triple selection module, a subgraph extraction module, a mask learning and interpretation path generation module.
[0081] The data acquisition and heterogeneous graph construction module is used to acquire drug synergy prediction related data and construct a multi-node type and multi-edge type bioinformatics heterogeneous graph based on the acquired drug synergy prediction related data; the graph representation learning module is used to learn the embedding features of nodes through the relational graph convolutional network, and aggregate the learned drug combination and cell line node embedding features to obtain the final embedding representation; the model training module is used to input the obtained final embedding representation into the multi-layer perceptron to output the probability of synergy between drug pairs on a specific cell line, thereby realizing the construction and training of the drug synergy prediction model; the synergy triplet The selection module is used to use the trained drug synergy prediction model to select synergistic drug, drug, and cell line triplets from the independent test set; the subgraph extraction module is used to use the nodes in each synergistic triple as the target node, and use the N-hop and K-core algorithms to calculate the biological information heterogeneous graph to obtain the subgraph; the mask learning and explanation path generation module is used to implement mask learning on all edges of the subgraph to filter out non-important edges, and use the Dijkstra algorithm in the subgraph after mask learning to find the path with the highest preset number of scores as the explanation for drug synergy prediction.
[0082] The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices. The terms "first", "second" and "third" are used to distinguish different objects, etc., and do not represent a sequence, nor do they limit the "first", "second" and "third" to different types.
[0083] In the description of the embodiments of this application, the words "exemplary," "for example," or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary," "for example," or "for example" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary," "for example," or "for example" is intended to present the relevant concepts in a concrete manner.
[0084] In the description of the embodiments of the present application, unless otherwise specified, “ / ” means or, for example, A / B can mean A or B; “and / or” in the text is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, “multiple” refers to two or more than two.
[0085] In some processes described in the embodiments of the present application, multiple operations or steps are included that appear in a specific order. However, it should be understood that these operations or steps may not be performed in the order in which they appear in the embodiments of the present application or may be performed in parallel. The sequence numbers of the operations are only used to distinguish between different operations, and the sequence numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations or steps may be performed in sequence or in parallel, and these operations or steps may be combined.
[0086] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above and includes a number of instructions for enabling a terminal device to execute the methods described in each embodiment of this application.
[0087] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A path-based graphical model interpretable method for explaining drug synergy prediction, characterized by: The path-based graph model interpretable method includes: Acquire drug synergy prediction related data, and construct a multi-node type and multi-edge type bioinformatics heterogeneous graph based on the acquired drug synergy prediction related data; The node embedding features are learned through a relational graph convolutional network, and the learned drug combination and cell line node embedding features are aggregated to obtain the final embedding representation; Inputting the obtained final embedding representation into a multi-layer perceptron to output the probability of synergy between drug pairs on a specific cell line, thereby realizing the construction and training of a drug synergy prediction model; Based on the trained drug synergy prediction model, synergistic drug, drug, and cell line triplets were selected from the independent test set; For each collaborative triple, the nodes in the triple are taken as target nodes, and the N-hop and K-core algorithms are used to calculate the biological information heterogeneous graph to obtain the subgraph; Mask learning is implemented on all edges of the subgraph to filter out non-important edges, and the Dijkstra algorithm is used in the subgraph after mask learning to find the path with the highest preset number of scores as the explanation for drug synergy prediction.
2. The path-based graphical model interpretability method according to claim 1, characterized in that: The drug synergy prediction related data includes drug-gene association data, cell line-gene association data, pathway-gene association data, and drug synergy data; The drug-gene association data are obtained based on FDA-approved or clinically investigated drugs; The cell line-gene association data were obtained from the Cancer Cell Line Encyclopedia; The pathway-gene association data were obtained from the therapeutic target database TTD; The drug synergy data is a drug combination dataset obtained from the DrugCombDB database. The drug combination dataset contains a variety of different drug combinations, and a virtual node is set for each drug combination to represent the drug-drug combination. The drugs in the drug combination are associated with the virtual nodes through directed edges.
3. The path-based graphical model interpretability method according to claim 2, characterized in that: The biological information heterogeneous graph includes multiple types of nodes and multiple types of edges; The nodes of the biological information heterogeneous graph include drugs, genes, pathways, cell lines, and virtual nodes; The edges of the biological information heterogeneous graph include virtual nodes-drug, drug-gene, pathway-gene, and cell line-gene; When processing different types of nodes and edges in biological information heterogeneous graphs, all types of nodes and edges are mapped to the same feature space. Biological entity nodes Represented as a vector , No. Relationship Type Represented as a vector ,in, Indicates the number of biological entity node types, represents the number of edge types, represents the embedding size of the biological entity node, Represents the embedding dimension of the relation type.
4. The path-based graphical model interpretable method according to claim 3, wherein: The embedding features of nodes learned through the relational graph convolutional network specifically include: In the bioinformatics heterogeneous graph R-GCN is used to learn the embedding features of nodes, where Represents a collection of nodes, represents the edge set, Represents a set of node types, Represents a set of edge types; In the In the R-GCN layer, for node Intermediate embedding The update is calculated as follows: in, represents a nonlinear activation function, Representation node In the layer neural network representation, Representation node In the layer neural network representation, Indicates The neighbor set of the node under the edge type, Used to balance the influence of neighbor nodes under different relationships, Representing relationships In the The learnable weight matrix of the layer, Representation node In the first The trainable self-loop weight matrix of the layer, Representation node In the Representation of layer neural network, nodes is a node The neighboring node.
5. The path-based graphical model interpretable method according to claim 4, characterized in that: The final embedding representation is input into a multilayer perceptron to output the probability of synergy between drug pairs on a specific cell line, specifically including: For the final embedding representation , which are connected in series to form a triplet Indicates that, 、 The final embedding features used to represent the two drugs, represents the final embedded features of a cell line, 、 Used to indicate two drugs, indicates cell lines; The triple The feature representation of the input multi-layer perceptron is used to calculate the score to determine the triple Is there a synergistic effect between them? The score function is expressed as: in, A triple representing the input The probability score of collaboration between represents the activation function, represents a multilayer perceptron, ‖ represents a concatenation operator; Binary cross entropy is used as the loss function to measure the difference between the predicted results and the actual results. Specifically: in, Represents the difference between the predicted result and the actual result, Represents a triple The total number of Represents a triple The true collaborative binary labels between them.
6. The path-based graphical model interpretable method according to claim 5, characterized in that: The drug synergy prediction model after training is used to select synergistic drug, drug, and cell line triplets from an independent test set, specifically including: Based on the trained drug synergy prediction model, the drug, drug, and cell line triplets in the independent test set are predicted to obtain the predicted synergy probability of the triplets; A threshold for binary classification is set. When the predicted synergistic probability of a triple is greater than the threshold, the current triple is judged to be a synergistic sample. When the predicted synergistic probability of a triple is not greater than the threshold, the current triple is judged to be a non-synergistic sample, thereby selecting synergistic triplets of drugs, drugs, and cell lines from the independent test set.
7. The path-based graphical model interpretable method according to claim 6, characterized in that: For each collaborative triple, the node in the triple is used as the target node, and the N-hop and K-core algorithms are used to calculate the biological information heterogeneous graph to obtain a subgraph, specifically including: The three nodes in the synergistic triplet and the virtual nodes corresponding to the drug combination are defined as seed nodes, and the graph is obtained after N-hop subgraph extraction. , the figure is the union of four subgraphs generated by four seed nodes; The K-core of a biological information heterogeneous graph is defined as the node with the minimum degree. The unique largest subgraph of the biological information heterogeneous graph is K-core, which is obtained by recursively removing the graph with a degree less than , until all nodes in the biological information heterogeneous graph have a degree of at least , in order to achieve the purpose of pruning, and set the nodes in the triplet not to participate in pruning; In the figure Use K-core pruning to obtain the final computational subgraph , thereby extracting a computational subgraph for each collaborative triple .
8. The path-based graphical model interpretable method according to claim 7, characterized in that: Mask learning is performed on all edges of the subgraph to filter out non-important edges, specifically including: set up Indicates the relationship type is All edges of Is the relationship type The learnable mask of all edges of , then the relationship type is Applying a mask on the edge set of Represented as , then the edge set that implements the mask on all edges of all relation types on the subgraph is represented as: in, represents the set of edges that enforces a mask on all edges of all relation types on the subgraph, represents the edge, represents the learned mask, represents element-wise product, Represents a set All relationship types in Perform a union operation; Using loss function and , measure the edges that have an impact on the output results of the drug synergy prediction model and the edges that can form meaningful paths; Among them, for the loss function , specifically: in, Represents minimizing the prediction loss between the original image and the mask image, Represents the drug synergy prediction model for triples on the original graph The predictions are collaborative, represents the mask image, represents the probability of predicting a triplet as a synergistic combination on the mask map; Among them, for the loss function , specifically: in, represents maximizing the weight of the edges on the explanation path, 、 represents the hyperparameter that balances the contributions of the two edges, represents a learnable mask for all edges of all types, represents a candidate edge, represents a non-candidate edge, Indicates type The edge, express learnable masks of type edges, Indicates type The learnable mask of the edges corresponds to the edge subset The weight value of .
9. The path-based graphical model interpretable method according to claim 8, wherein: The method of using the Dijkstra algorithm to find the preset number of pathways with the highest scores as the explanation for drug synergy prediction specifically includes: Combined Edges The ability to influence the prediction results and the degree of the target node, calculate the edge Importance score: in, Represents an edge The importance score of Represents an edge The degree of the target node in , Represents an edge The probability of being selected as the explanatory path, Represents an edge The mask weight of In a biological information heterogeneous graph, a path is formed by a set of edges, and the score of the path is equivalent to the sum of the scores of each edge included in the path. Specifically: in, represents the path score, Indicates the path, Represents an edge The probability of being included in the explanatory path, ; Based on the calculated path score, combined with the shortest path algorithm Dijistra, and using the negative score of the edge As the score of each edge, a preset number of top-ranked shortest paths are obtained as target paths, thereby obtaining an explanation for drug synergy prediction.
10. A path-based graphical model interpretable device, characterized in that: The path-based graph model interpretable device includes: A data acquisition and heterogeneous graph construction module is used to acquire drug synergy prediction related data and construct a multi-node and multi-edge bioinformatics heterogeneous graph based on the acquired drug synergy prediction related data; A graph representation learning module, which is used to learn node embedding features through a relational graph convolutional network and aggregate the learned drug combination and cell line node embedding features to obtain the final embedding representation; a model training module, which is used to input the obtained final embedding representation into a multi-layer perceptron to output the probability of synergy between drug pairs on a specific cell line, thereby realizing the construction and training of a drug synergy prediction model; A synergistic triplet selection module is used to select synergistic drug, drug, and cell line triplets from an independent test set using a trained drug synergy prediction model; The subgraph extraction module is used to extract the subgraph from each collaborative triplet, taking the nodes in the triplet as the target nodes and applying the N-hop and K-core algorithms to the biological information heterogeneous graph; The mask learning and explanation path generation module is used to implement mask learning on all edges of the subgraph to filter out non-important edges, and use the Dijkstra algorithm in the subgraph after mask learning to find the path with the highest preset number of scores as the explanation for drug synergy prediction.