SHAP Data Search Compression and Similarity Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for searching SHAP values, which represent the influence degree of features in AI models, face challenges in achieving both speed and accuracy due to the large amount of data involved.

Innovation Solution

A search support device and method that calculates and compresses SHAP data, allowing for the generation of compressed SHAP matrices, and then calculates similarities between these matrices to identify features that meet predetermined conditions, thereby facilitating high-speed and accurate searches.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If SHAP data is stored in full detail to maintain accuracy, then measurement precision is improved, but productivity deteriorates due to enormous data amount slowing down search operations

Engineering Contradiction:
Improveaccuracy of SHAP value searchVSAvoidspeed of SHAP value search
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments SHAP data into multiple dimensions including feature name, feature type, SHAP value range, and data type. By dividing the enormous SHAP dataset into these manageable segments, the system can perform targeted searches without processing the entire dataset, thus improving search speed while maintaining accuracy through multi-dimensional filtering capabilities.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces multiple parameters for organizing and searching SHAP data, including feature type classifications (numeric, categorical, text), SHAP value ranges (positive, negative, zero), and data type categories. These parameter changes enable efficient filtering and searching of SHAP data without requiring full data processing, resolving the contradiction between search accuracy and speed.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If comprehensive SHAP data features are maintained for various applications, then adaptability is improved, but device complexity worsens due to the need to handle enormous and diverse data

Engineering Contradiction:
Improveapplicability of SHAP dataVSAvoidcomplexity of data processing system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal SHAP data processing system that handles multiple data types (numeric, categorical, text) and feature types through a unified framework. The search device can process diverse SHAP data for various applications using the same multi-dimensional organization structure, reducing system complexity while maintaining high adaptability across different use cases.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent adds multiple organizational dimensions to SHAP data including feature type, SHAP value range, and data type classifications. By organizing data along these additional dimensions, the system can efficiently handle diverse SHAP data for various applications without increasing processing complexity, as the multi-dimensional structure provides systematic access paths for different query types.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20230325692A1Search support device and search support method
Publication Date: 2023.10.12 HITACHI LTD
  • US20230325692A1 patent drawing
  • US20230325692A1 patent drawing
  • US20230325692A1 patent drawing

AI summary

Provided is a search support device to perform a search related to a parameter representing an influence degree of a feature at high speed and with high accuracy. The search support device calculates at least one or more pieces of SHAP data indicating an influence degree of each feature in a trained model on output data output from the trained model; a process of generating compressed SHAP data, which is data obtained by compressing the SHAP data, for each of the SHAP data and storing the compressed SHAP data; verification target SHAP data which is SHAP data for output data output from the trained model by inputting input data to the trained model; a similarity between each of the calculated compressed SHAP data and the calculated verification target SHAP data, and specifies the compressed SHAP data in which the similarity with the verification target SHAP data satisfies a predetermined condition.