Graph Data Neighborhood Generation for Explainable ML
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data generation techniques, such as LIME, struggle to effectively generate neighborhood data for machine learning models using graph data, often resulting in impaired features that hinder the application of these models in mission-critical areas where accountability is required.
Innovation Solution
A data generation program and device that modifies the connection relationships in graph data by changing selected edges within a threshold, maintaining connectivity and structure, to generate suitable neighborhood data for linear regression models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional data generation techniques (e.g., LIME) are used to generate neighborhood data for graph-based machine learning models, then the data can be generated with varied feature amounts, but the generated data often has impaired features that reduce its quality and reliability
Solution Approach 1:
The patent applies local quality by selectively modifying only specific edges in the graph data while preserving the overall structure. Instead of randomly varying all features, the invention identifies and modifies only the necessary local connections to create meaningful neighborhood data variations that maintain graph integrity and feature quality.
Solution Approach 2:
The patent segments the graph data into nodes and edges, allowing independent manipulation of connection relationships. By separating the modification process into node selection and edge connection changes, the invention can generate diverse neighborhood data while maintaining structural coherence and avoiding feature impairment.
2Loss of information
If the connection relationships in graph data are modified to generate diverse neighborhood data, then the explainability of machine learning models can be improved, but the structural integrity and connectivity of the graph data may be compromised
Solution Approach 1:
The patent applies preliminary action by first selecting nodes and determining their connection relationships before actually modifying the edges. This preliminary planning ensures that the modified graph maintains structural integrity and connectivity, preventing the loss of important graph properties while still achieving diverse neighborhood data for improved model explainability.
Solution Approach 2:
The patent uses the node selection process as an intermediary between the original graph structure and the modified neighborhood data. By introducing this intermediate step where nodes are selected and their connections are planned before modification, the invention bridges the gap between maintaining structural integrity and achieving data diversity for better explainability.
Data Source
AI summary
A non-transitory computer-readable storage medium storing a data generation program for causing a computer to perform processing including: obtaining data that includes a plurality of nodes and a plurality of edges connecting the plurality of nodes; selecting a first edge from the plurality of edges; and generating new data that has a second connection relationship between the plurality of nodes different from a first connection relationship between the plurality of nodes of the data by changing connection of the first edge such that a third node connected to at least one of a first node and a second node located at both ends of the first edge via a number of edges, the number being equal to or less than a threshold, is located at one end of the first edge.


