Graph Data Neighborhood Generation for Explainable ML

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data generation techniques, such as LIME, struggle to effectively generate neighborhood data for machine learning models using graph data, often resulting in impaired features that hinder the application of these models in mission-critical areas where accountability is required.

Innovation Solution

A data generation program and device that modifies the connection relationships in graph data by changing selected edges within a threshold, maintaining connectivity and structure, to generate suitable neighborhood data for linear regression models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional data generation techniques (e.g., LIME) are used to generate neighborhood data for graph-based machine learning models, then the data can be generated with varied feature amounts, but the generated data often has impaired features that reduce its quality and reliability

Engineering Contradiction:
Improveadaptability to graph dataVSAvoidquality of generated neighborhood data
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent applies local quality by selectively modifying only specific edges in the graph data while preserving the overall structure. Instead of randomly varying all features, the invention identifies and modifies only the necessary local connections to create meaningful neighborhood data variations that maintain graph integrity and feature quality.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the graph data into nodes and edges, allowing independent manipulation of connection relationships. By separating the modification process into node selection and edge connection changes, the invention can generate diverse neighborhood data while maintaining structural coherence and avoiding feature impairment.

Inventive Principle:
Principle #1Segmentation

2Loss of information

If the connection relationships in graph data are modified to generate diverse neighborhood data, then the explainability of machine learning models can be improved, but the structural integrity and connectivity of the graph data may be compromised

Engineering Contradiction:
Improveexplainability of classification resultsVSAvoidstructural integrity of graph data
Core Design Contradiction:
Loss of informationVSStability of the object's composition

Solution Approach 1:

The patent applies preliminary action by first selecting nodes and determining their connection relationships before actually modifying the edges. This preliminary planning ensures that the modified graph maintains structural integrity and connectivity, preventing the loss of important graph properties while still achieving diverse neighborhood data for improved model explainability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses the node selection process as an intermediary between the original graph structure and the modified neighborhood data. By introducing this intermediate step where nodes are selected and their connections are planned before modification, the invention bridges the gap between maintaining structural integrity and achieving data diversity for better explainability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20230196129A1Non-transitory computer-readable storage medium for storing data generation program, data generation method, and data generation device
Publication Date: 2023.06.22 FUJITSU LTD
  • US20230196129A1 patent drawing
  • US20230196129A1 patent drawing
  • US20230196129A1 patent drawing

AI summary

A non-transitory computer-readable storage medium storing a data generation program for causing a computer to perform processing including: obtaining data that includes a plurality of nodes and a plurality of edges connecting the plurality of nodes; selecting a first edge from the plurality of edges; and generating new data that has a second connection relationship between the plurality of nodes different from a first connection relationship between the plurality of nodes of the data by changing connection of the first edge such that a third node connected to at least one of a first node and a second node located at both ends of the first edge via a number of edges, the number being equal to or less than a threshold, is located at one end of the first edge.