Automated Feature Value Generation for Data Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data analysis systems struggle to automatically create feature values that are highly correlated with objective indices, especially as the size of the data increases, making it difficult for analysts to manually process and detect relevance in large datasets.

Innovation Solution

A data analysis support system that includes a processor and storage device configured to select explanatory index items, perform clustering, determine value ranges, and generate feature values that can be easily interpreted by humans, allowing for automated creation of feature values from input tables.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual data processing is used to detect relevance, then analysis accuracy can be maintained, but productivity decreases as data size increases

Engineering Contradiction:
Improverelevance detection accuracyVSAvoiddata processing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system enables automated self-service data analysis by having the computer automatically perform clustering analysis on explanatory indices and generate feature values without requiring manual analyst intervention. The system processes large datasets autonomously, identifying correlations and generating analysis results that would otherwise require manual examination.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual processing system with an automated computational system. Instead of analysts manually examining data relationships, the system uses computer-based clustering algorithms and automated feature value generation to detect relevance and correlations in large datasets.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Device complexity

If fixed columns are used as feature values, then device complexity is reduced, but adaptability decreases for creating highly correlated feature values

Engineering Contradiction:
Improvesystem simplicityVSAvoidfeature value correlation capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The system transitions from static fixed columns to dynamic feature value generation. The clustering analysis dynamically identifies relationships in the data, and feature values are automatically generated based on these identified relationships, allowing the system to adapt to different datasets and correlation patterns rather than being constrained to predetermined columns.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the parameters used for analysis by automatically generating feature values based on clustering results rather than using fixed column values. This parameter transformation allows the system to create highly correlated feature values that are specifically tailored to the relationships present in each dataset.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10509808B2Data analysis support system and data analysis support method
Publication Date: 2019.12.17 HITACHI LTD
  • US10509808B2 patent drawing
  • US10509808B2 patent drawing
  • US10509808B2 patent drawing

AI summary

Provided is a data analysis support system, comprising a processor, and a storage device which is coupled to the processor. The storage device retains objective index information which associates primary key values with objective index values, and explanatory index information which associates values common to the primary keys with sets of values of explanatory indices of a plurality of items. The processor selects one or more items of the explanatory indices, clusters the values of the explanatory indices of the selected one or more items, identifies a range of the values of the explanatory indices of each of the items of each of the clusters which are obtained by the clustering, and outputs the identified value ranges.