Partial Graph Structure Selection Program, Apparatus, and Method
By calculating and ranking subgraph structures based on frequency, standard deviation, and explanatory score, the method addresses the challenge of selecting significant subgraph structures for graph kernels, enhancing model accuracy and explanation conciseness.
Patent Information
- Application Number
- JP2024510951
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-30
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2042-03-30
AI Technical Summary
Existing graph XAI technologies face challenges in selecting significant subgraph structures for graph kernels, as methods that delete graphlets with low frequency or redundancy can inadvertently exclude important subgraphs, affecting the accuracy of machine learning models.
The proposed solution calculates the occurrence frequency and explanatory score of each subgraph structure, then uses a product of average frequency, standard deviation, and explanatory score to rank and select subgraph structures. This process iteratively adds and evaluates subgraph structures based on model accuracy, ensuring significant structures are retained.
This approach effectively selects significant subgraph structures for graph kernels, improving the accuracy of machine learning models while reducing computational costs, and enabling concise and meaningful explanations of prediction results.
Smart Images

Figure 0007694809000001 
Figure 0007694809000002 
Figure 0007694809000003
Abstract
Description
Technical Field
[0001] The disclosed technology relates to a sub-graph structure selection program, a sub-graph structure selection device, and a sub-graph structure selection method.
Background Art
[0002] There is a technology that inputs graph data into a machine learning model such as a neural network, obtains a prediction result according to a task, and obtains a sub-graph that contributed to the prediction result from the input graph data. This sub-graph is information that can explain the prediction process by the machine learning model. Also, a machine learning model that can explain the prediction process in this way is called explainable AI (XAI: Explainable Artificial Intelligence), and XAI whose input data is graph data is called graph XAI.
[0003] In order to use graph data as an input to a machine learning model, there is a technology called a graph kernel that maps graph data to a high-dimensional vector. Examples of graph kernels include Random walk kernel, Graphlet kernel, Weisfeiler-Lehman kernel, etc. In many cases, each element of the mapped vector of these graph kernels represents a primitive sub-graph. In graph XAI, it is desirable to obtain a vector representation of graph data that is as concise as possible.
[0004] For example, the Graphlet kernel enumerates graphlets consisting of a small number of nodes, counts the number of times each graphlet appears in the graph, and vectorizes the graph. A graphlet contains a predetermined number of nodes and enumerates all patterns of connections between the nodes. When the number of nodes is {3, 4, 5}, the number of graphlets is 29, so the vector is 29-dimensional. In the vectorization of the graph using this graphlet, there is a problem that the computational cost of counting the graphlets is high. In order to reduce the computational cost, for example, it is conceivable to reduce the number of graphlets by restricting the number of nodes, such as setting the number of nodes of the graphlet to {3, 4}. However, in this case, simply reducing the number of graphlets is not possible because it has an adverse effect on the accuracy of learning and prediction of the machine learning model using the graph data vector.
[0005] Therefore, a technique for selecting graphlets has been proposed to reduce the computational cost of graph vectorization and improve the accuracy of learning and prediction. This technique focuses on the fact that in a graph of a specific domain, the appearance frequency of a specific graphlet is often low, and deletes graphlets with a low appearance frequency or standard deviation in the graph. In addition, this technique deletes redundant graphlets that are highly correlated with other graphlets.
Prior Art Documents
Non-Patent Documents
[0006]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0007] However, in the above prior art, graphlets corresponding to significant subgraphs in the graph may be deleted from the graphlets due to low frequency or redundancy.
[0008] As one aspect, the disclosed technology aims to select a significant subgraph structure as the subgraph structure to be used in the graph kernel.
Means for Solving the Problem
[0009] As one embodiment, the disclosed technology calculates the occurrence frequency of each of a plurality of predetermined subgraph structures in each of one or more graphs to be predicted, the graphs including a plurality of nodes and a plurality of edges. Further, the disclosed technology calculates the explanatory score of each of the plurality of subgraph structures based on the contribution degree of each of the nodes or the edges to the prediction result output when each of the one or more graphs to be predicted is input to a trained machine learning model. Further, the disclosed technology calculates, for each of the plurality of subgraph structures, the product of the average of the occurrence frequencies, the standard deviation of the occurrence frequencies, and the average of the explanatory scores in the one or more graphs to be predicted. Further, the disclosed technology selects one subgraph structure from among the plurality of subgraph structures in descending order of the product and adds it to a list. Each time, the disclosed technology calculates the accuracy of the machine learning model when the vectorized graph to be predicted using the subgraph structures included in the list is input. Then, when the change in the accuracy satisfies a predetermined condition, the disclosed technology selects the subgraph structure added to the list as the subgraph structure to be finally used.
Advantages of the Invention
[0010] As one aspect, it has the effect that a significant subgraph structure can be selected as the subgraph structure to be used in the graph kernel.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Embodiments for Carrying Out the Invention
[0012] Hereinafter, an example of an embodiment according to the disclosed technology will be described with reference to the drawings.
[0013] As shown in FIG. 1, an explanatory graph set is input to the partial graph structure selection device 10. The partial graph structure selection device 10 selects and outputs graphlets to be used in a graph kernel based on the explanatory graph set. Note that a graphlet is an example of the "partial graph structure" of the disclosed technology.
[0014] An explanatory graph is a graph including a plurality of nodes and a plurality of edges connecting the nodes, and is a graph in which a contribution degree to a prediction result output when input to a trained machine learning model, that is, a contribution degree to the prediction, is assigned to each node or edge. In the present embodiment, the case where the contribution degree is assigned to each node will be described as an example.
[0015] In the upper part of FIG. 2, an example of an explanatory graph is shown. In the upper part of FIG. 2, an example of a graph representing a chemical structure is shown. The numbers written next to each node (circle) are the contribution degrees. In the present embodiment, under the assumption that the average of the contribution degrees of the nodes included in the significant part with respect to the prediction result becomes high, the contribution degree is used for the selection of graphlets. In the case of a chemical structure as shown in FIG. 2, this assumption corresponds to the fact that the average of the contribution degrees of the nodes constituting the chemically significant structure becomes high. For example, in the explanatory graph in the upper part of FIG. 2, between the partial graph shown in A in the lower part of FIG. 2 (hereinafter referred to as "partial graph A") and the partial graph shown in B (hereinafter referred to as "partial graph B"), the average of the contribution degrees is higher in partial graph B. Therefore, partial graph B represents a significant structure in the explanatory graph.
[0016] Here, when selecting graphlets to be used in the graph kernel as in the above-described prior art, graphlets with low appearance frequencies or standard deviations in the explanatory graph are deleted, and redundant graphlets with high correlation with other graphlets are deleted. In the example of FIG. 2 above, when partial graph A and partial graph B appear simultaneously with high frequency, as shown in FIG. 3, for the reason of redundancy, there is a possibility that the graphlet whose structure matches the partial graph representing the chemically meaningful structure CH3 is deleted. That is, there is a possibility that significant graphlets are excluded from the graphlets to be used in the graph kernel. Therefore, in the present embodiment, as described above, the contribution degree assigned to each node of the explanatory graph is used for the selection of graphlets.
[0017] Functionally, the partial graph structure selection device 10 includes an appearance frequency calculation unit 12, an explanation score calculation unit 14, an evaluation value calculation unit 16, a selection unit 18, and a deletion unit 20. Also, a prediction model 30, which is a trained machine learning model, is stored in a predetermined storage area of the partial graph structure selection device 10. Note that the evaluation value calculation unit 16 is an example of the "product calculation unit" of the disclosed technology.
[0018] The occurrence frequency calculation unit 12 calculates the occurrence frequency of each of a plurality of predetermined graphlets in each of the explanatory graphs included in the set of explanatory graphs. As the plurality of predetermined graphlets, for example, as shown in FIG. 4, graphlets g1 to g with the number of nodes being {3, 4, 5} 29 may be defined as 29 graphlets. The diagrams of the graphlets used in a part of FIG. 4 and FIGS. 5 and 8 described later are cited from the drawings of Non-Patent Document 1. The occurrence frequency calculation unit 12 calculates the occurrence frequency, for example, as shown in FIG. 5, by searching for and counting sub-graphs in the explanatory graph whose structure matches that of each graphlet (in the example of FIG. 5, graphlet g6). In the example of FIG. 5, as sub-graphs whose structure matches that of graphlet g6, there are sub-graph A (the sub-graph indicated by the broken line) and sub-graph B (the sub-graph indicated by the one-dot chain line). Therefore, the occurrence frequency calculation unit 12 calculates the occurrence frequency of graphlet g6 as "2".
[0019] The explanatory score calculation unit 14 calculates the explanatory score of each graphlet based on the contribution degree of each node of the explanatory graph. Specifically, the explanatory score calculation unit 14 calculates, in the explanatory graph, the average of the contribution degrees of the nodes included in the sub-graph whose structure matches that of the graphlet as the explanatory score of that graphlet.
[0020] When there are a plurality of sub-graphs in one explanatory graph whose structures match that of the graphlet, the explanatory score calculation unit 14 sets the higher of the explanatory scores calculated for each of the plurality of sub-graphs as the explanatory score of that graphlet. In the case of the example of FIG. 5, the explanatory score of sub-graph A is (0.2 + 0.6 + 0.9 + 0.1) / 4 = 0.45, and the explanatory score of sub-graph B is (0.8 + 0.9 + 0.7 + 0.7) / 4 = 0.775. Therefore, the explanatory score calculation unit 14 calculates the explanatory score of graphlet g6 for the corresponding explanatory graph as 0.775.
[0021] Note that the explanation score calculation unit 14 is not limited to selecting the sub-graph with the higher explanation score when there are multiple sub-graphs that match the structure of the graphlet. Instead, the average of the explanation scores for the multiple sub-graphs may be calculated as the explanation score for the corresponding graphlet.
[0022] The evaluation value calculation unit 16 calculates, for each of the multiple graphlets, the product of the average of the appearance frequencies, the standard deviation of the appearance frequencies, and the average of the explanation scores in the set of explanatory graphs as the evaluation value. Specifically, the evaluation value calculation unit 16 calculates, for graphlet g i the average (hereinafter referred to as "average appearance frequency") μ i of the appearance frequencies calculated from each explanatory graph over all explanatory graphs. Also, the evaluation value calculation unit 16 calculates, for graphlet g i the standard deviation σ i of the appearance frequencies calculated from each explanatory graph over all explanatory graphs. Further, the evaluation value calculation unit 16 calculates, for graphlet g i the average (hereinafter referred to as "average explanation score") s i of the explanation scores calculated from each explanatory graph over all explanatory graphs. Then, the evaluation value calculation unit 16 calculates the product of the average appearance frequency μ i the standard deviation σ i and the average explanation score s i as the evaluation value μσs i of graphlet g i
[0023] The selection unit 18 selects one graplet from among a plurality of graplets in descending order of the evaluation value calculated by the evaluation value calculation unit 16 and adds it to the list. Each time, the selection unit 18 calculates the accuracy of the prediction model 30 when using the explanatory graph vectorized using the graplets included in the list as input. When the change in accuracy satisfies a predetermined condition, the selection unit 18 passes the list to the deletion unit 20. The selection unit 18 may set the predetermined condition as the case where the accuracy no longer increases or decreases. The selection unit 18 may determine that the case where the difference between the previously calculated accuracy and the currently calculated accuracy is within a predetermined value is the case where the accuracy no longer increases. Also, the selection unit 18 may determine that the case where the currently calculated accuracy is lower than the previously calculated accuracy is the case where the accuracy decreases.
[0024] The deletion unit 20 calculates an index indicating the correlation of all pairs of the graplets added to the list, and deletes from the list the graplet with the lower average explanation score s for the pairs with the index being equal to or greater than a predetermined value. The deletion unit 20 may calculate the cross-correlation c as the index indicating the correlation. If both of the highly correlated graplets are left, it will be redundant, so one of them is deleted. At that time, by deleting the graplet with the lower average explanation score s, it becomes easier for the graplets with significant structures to remain. The deletion unit 20 finally outputs the graplets remaining in the list as the graplets to be used by the graph kernel.
[0025] The partial graph structure selection device 10 may be implemented by, for example, a computer 40 shown in FIG. 6. The computer 40 includes a CPU (Central Processing Unit) 41, a memory 42 as a temporary storage area, and a non-volatile storage unit 43. The computer 40 also includes input / output devices 44 such as an input unit and a display unit, and an R / W (Read / Write) unit 45 that controls reading and writing of data to and from a storage medium 49. The computer 40 further includes a communication I / F (Interface) 46 connected to a network such as the Internet. The CPU 41, the memory 42, the storage unit 43, the input / output devices 44, the R / W unit 45, and the communication I / F 46 are connected to each other via a bus 47.
[0026] The storage unit 43 may be implemented by an HDD (Hard Disk Drive), an SSD (Solid State Drive), a flash memory, or the like. In the storage unit 43 as a storage medium, a partial graph structure selection program 50 for causing the computer 40 to function as the partial graph structure selection device 10 is stored. The partial graph structure selection program 50 includes an appearance frequency calculation process 52, an explanation score calculation process 54, an evaluation value calculation process 56, a selection process 58, and a deletion process 60. The storage unit 43 also has an information storage area 70 in which information constituting the prediction model 30 is stored.
[0027] The CPU 41 reads the partial graph structure selection program 50 from the storage unit 43, expands it in the memory 42, and sequentially executes the processes included in the partial graph structure selection program 50. By executing the occurrence frequency calculation process 52, the CPU 41 operates as the occurrence frequency calculation unit 12 shown in FIG. 1. Also, by executing the explanation score calculation process 54, the CPU 41 operates as the explanation score calculation unit 14 shown in FIG. 1. Further, by executing the evaluation value calculation process 56, the CPU 41 operates as the evaluation value calculation unit 16 shown in FIG. 1. Additionally, by executing the selection process 58, the CPU 41 operates as the selection unit 18 shown in FIG. 1. Moreover, by executing the deletion process 60, the CPU 41 operates as the deletion unit 20 shown in FIG. 1. Also, the CPU 41 reads information from the information storage area 70 and expands the prediction model 30 in the memory 42. As a result, the computer 40 that has executed the partial graph structure selection program 50 functions as the partial graph structure selection device 10. Note that the CPU 41 that executes the program is hardware.
[0028] Note that the functions realized by the partial graph structure selection program 50 can also be realized by, for example, a semiconductor integrated circuit, and more specifically, an ASIC (Application Specific Integrated Circuit) or the like.
[0029] Next, the operation of the partial graph structure selection device 10 according to the present embodiment will be described. When an explanation graph set is input to the partial graph structure selection device 10 and the selection of graphlets is instructed, the partial graph structure selection process shown in FIG. 7 is executed in the partial graph structure selection device 10. Note that the partial graph structure selection process is an example of the partial graph structure selection method of the disclosed technology.
[0030] In step S10, the occurrence frequency calculation unit 12 acquires the explanation graph set input to the partial graph structure selection device 10. Next, in step S12, the occurrence frequency calculation unit 12 searches for subgraphs in the explanation graph whose graphlets match the structure and counts them, thereby calculating the occurrence frequency of each graphlet in each explanation graph.
[0031] Next, in step S14, the explanation score calculation unit 14 calculates, as the explanation score of the graphlet, the average of the contribution degrees of the nodes included in the sub-graph that matches the structure of the graphlet in the explanation graph. The explanation score calculation unit 14 calculates the explanation score of each graphlet in each explanation graph.
[0032] Next, in step S16, the evaluation value calculation unit 16 calculates, for each graphlet, the average appearance frequency that is the average of the appearance frequencies calculated from each explanation graph, the standard deviation of the appearance frequencies, and the average explanation score that is the average of the explanation scores calculated from each explanation graph. Then, the evaluation value calculation unit 16 calculates the product of the average appearance frequency, the standard deviation, and the average explanation score as the evaluation value of each graphlet.
[0033] Next, in step S18, the selection unit 18 creates a list L in which a plurality of graphlets are sorted in descending order of the evaluation value calculated in step S16 above. Next, in step S20, the selection unit 18 selects the graphlet with the maximum evaluation value from the list L, adds it to the list L', and deletes it from the list L.
[0034] Next, in step S22, the selection unit 18 calculates the accuracy of the prediction model 30 when using the graphlet included in the list L' as a graph kernel and inputting the vectorized explanation graph. Next, in step S24, the selection unit 18 determines whether or not the accuracy calculated in step S22 above has decreased compared to the accuracy calculated last time. If the accuracy has not decreased, the process returns to step S20, and if it has decreased, the process proceeds to step S26.
[0035] In step S26, the selection unit 18 deletes the graplet that was last added to the list L' from the list L', and passes the list L' to the deletion unit 20. Next, in step S28, the deletion unit 20 calculates an index indicating the correlation of all pairs of graplets in the list L'. Then, for pairs with an index indicating correlation greater than or equal to a predetermined value, the deletion unit 20 deletes the graplet with the lower average explanation score s from the list L'. The deletion unit 20 finally outputs the graplets remaining in the list L' as the graplets to be used in the graph kernel, and ends the sub-graph structure selection process.
[0036] As described above, the sub-graph structure selection device according to the present embodiment calculates the appearance frequency of each of a plurality of predetermined graplets in each of one or more explanatory graphs including a plurality of nodes and a plurality of edges. Further, the sub-graph structure selection device calculates the explanation score of each of the plurality of graplets based on the contribution degree of each node given to the explanatory graph. Further, the sub-graph structure selection device calculates, for each of the plurality of sub-graplets, the average appearance frequency, the standard deviation of the appearance frequency, and the product of the average explanation score in the explanatory graph set as an evaluation value. Further, the sub-graph structure selection device selects one graplet from among the plurality of graplets in descending order of the evaluation value and adds it to the list. Each time, the sub-graph structure selection device calculates the accuracy of the prediction model when the explanatory graph vectorized using the graplets included in the list is input. Then, when the change in accuracy satisfies a predetermined condition, the sub-graph structure selection device selects the graplet added to the list as the sub-graph structure to be finally used in the graph kernel. Thereby, a significant sub-graph structure can be selected as the sub-graph structure to be used in the graph kernel.
[0037] For example, as shown in FIG. 8, in this embodiment, by selecting graphlets used for a graph kernel using an explanation score based on the contribution degree in an explanatory graph, it is possible to select a concise combination of graphlets without losing a significant partial graph structure. Then, by using the graphlets selected in this way as a graph kernel and performing prediction by a machine learning model on the graph to be predicted, the prediction result explanation obtained together with the prediction result can also be expressed in a concise combination without losing significance. In the example of FIG. 8, among the selected graphlets, a partial graph (thick line part) whose structure matches the graphlet surrounded by a broken line is shown as an example of a partial graph contributing to the prediction.
[0038] As a result, when performing subsequent causal inference or the like based on the prediction result and the prediction result explanation, it becomes easier to estimate a significant causal relationship as the causal relationship or the like between partial graphs in the graph. For example, in the case of a graph representing a chemical structure, performing causal inference can contribute to discovering a partial graph related to the reaction mechanism.
[0039] Note that in the above embodiment, an aspect in which the partial graph structure selection program is pre-stored (installed) in the storage unit has been described, but the present invention is not limited to this. The program according to the disclosed technology can also be provided in a form stored in a storage medium such as a CD-ROM, a DVD-ROM, or a USB memory.
Explanation of Signs
[0040] 10 Partial graph structure selection device 12 Appearance frequency calculation unit 14 Explanation score calculation unit 16 Evaluation value calculation unit 18 Selection unit 20 Deletion unit 30 Prediction model 40 Computer 41 CPU 42 Memory 43 Storage unit 44 Input / output device 45 R / W unit 46 Communication I / F 47 Bus 49 Memory Medium 50 Sub - graph Structure Selection Program 52 Appearance Frequency Calculation Process 54 Explanation Score Calculation Process 56 Evaluation Value Calculation Process 58 Selection Process 60 Deletion Process 70 Information Storage Area
Claims
1. calculating the occurrence frequency of each of a plurality of predetermined sub - graph structures in each of one or more graphs to be predicted, the graphs including a plurality of nodes and a plurality of edges; calculating an explanatory score for each of the plurality of sub - graph structures based on the contribution degree of each of the nodes or the edges to the prediction result output when each of the one or more graphs to be predicted is input into a trained machine learning model; for each of the plurality of sub - graph structures, calculating the product of the average of the occurrence frequencies, the standard deviation of the occurrence frequencies, and the average of the explanatory scores in the one or more graphs to be predicted; each time one sub - graph structure is selected from the plurality of sub - graph structures in descending order of the product and added to a list, calculating the accuracy of the machine learning model when the graph to be predicted vectorized using the sub - graph structures included in the list is used as an input, and when the change in the accuracy satisfies a predetermined condition, selecting the sub - graph structure added to the list as the finally used sub - graph structure A sub - graph structure selection program for causing a computer to execute a process including this.
2. calculating an index indicating the correlation of all pairs of the sub - graph structures added to the list, and deleting from the list the sub - graph structure with the lower explanatory score for pairs with the index greater than or equal to a predetermined value. The sub - graph structure selection program according to claim 1.
3. The sub - graph structure selection program according to claim 1 or claim 2, wherein as the explanatory score, the average of the contribution degrees of the nodes or the edges included in a sub - graph having the same structure as the sub - graph structure in one graph to be predicted is calculated.
4. If there are multiple sub-graphs in one of the graphs to be predicted that have the same structure as the sub-graph structure, the higher value among the averages of the contribution degrees calculated for each of the multiple sub-graphs is used as the explanation score of the sub-graph structure. The sub-graph structure selection program according to claim 3.
5. The predetermined condition is that the difference between the accuracy calculated last time and the accuracy calculated this time is within a predetermined value, or the accuracy calculated this time is lower than the accuracy calculated last time. The sub-graph structure selection program according to claim 1 or claim 2.
6. An appearance frequency calculation unit that calculates the appearance frequency of each of a plurality of predetermined sub-graph structures in each of one or more graphs to be predicted, including a plurality of nodes and a plurality of edges. An explanation score calculation unit that calculates the explanation score of each of the plurality of sub-graph structures based on the contribution degree of each node or edge to the prediction result output when each of the one or more graphs to be predicted is input to a trained machine learning model. A product calculation unit that calculates the product of the average of the appearance frequencies, the standard deviation of the appearance frequencies, and the average of the explanation scores in the one or more graphs to be predicted for each of the plurality of sub-graph structures. Each time one sub-graph structure is selected from the plurality of sub-graph structures in descending order of the product and added to the list, the accuracy of the machine learning model when the vectorized graph to be predicted using the sub-graph structures included in the list is input is calculated. When the change in the accuracy satisfies a predetermined condition, the sub-graph structure added to the list is selected as the finally used sub-graph structure. A selection unit. A sub-graph structure selection device including the above.
7. Calculate the appearance frequency of each of a plurality of predetermined sub-graph structures in each of one or more graphs to be predicted, including a plurality of nodes and a plurality of edges. Based on the contribution degree for each of the nodes or edges with respect to the prediction results output when each of the graphs of one or more prediction targets is input into a trained machine learning model, calculate the explanation score for each of the plurality of sub-graph structures. For each of the plurality of sub-graph structures, calculate the product of the average of the occurrence frequencies, the standard deviation of the occurrence frequencies, and the average of the explanation scores in the graph of the one or more prediction targets. Each time one sub-graph structure is selected from among the plurality of sub-graph structures in descending order of the product and added to the list, calculate the accuracy of the machine learning model when the vectorized graph of the prediction target using the sub-graph structures included in the list is used as the input. When the change in the accuracy satisfies a predetermined condition, select the sub-graph structure added to the list as the sub-graph structure to be finally used. A sub-graph structure selection method in which a computer executes a process including this.
Citation Information
Patent Citations
Explaining graph-based predictions using network motif analysis
JP2022117452A