Automated Feature Value Generation for Data Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data analysis systems struggle to automatically create feature values that are highly correlated with objective indices, especially as the size of the data increases, making it difficult for analysts to manually process and detect relevance in large datasets.
Innovation Solution
A data analysis support system that includes a processor and storage device configured to select explanatory index items, perform clustering, determine value ranges, and generate feature values that can be easily interpreted by humans, allowing for automated creation of feature values from input tables.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual data processing is used to detect relevance, then analysis accuracy can be maintained, but productivity decreases as data size increases
Solution Approach 1:
The system enables automated self-service data analysis by having the computer automatically perform clustering analysis on explanatory indices and generate feature values without requiring manual analyst intervention. The system processes large datasets autonomously, identifying correlations and generating analysis results that would otherwise require manual examination.
Solution Approach 2:
The patent replaces the mechanical manual processing system with an automated computational system. Instead of analysts manually examining data relationships, the system uses computer-based clustering algorithms and automated feature value generation to detect relevance and correlations in large datasets.
2Device complexity
If fixed columns are used as feature values, then device complexity is reduced, but adaptability decreases for creating highly correlated feature values
Solution Approach 1:
The system transitions from static fixed columns to dynamic feature value generation. The clustering analysis dynamically identifies relationships in the data, and feature values are automatically generated based on these identified relationships, allowing the system to adapt to different datasets and correlation patterns rather than being constrained to predetermined columns.
Solution Approach 2:
The system changes the parameters used for analysis by automatically generating feature values based on clustering results rather than using fixed column values. This parameter transformation allows the system to create highly correlated feature values that are specifically tailored to the relationships present in each dataset.
Data Source
AI summary
Provided is a data analysis support system, comprising a processor, and a storage device which is coupled to the processor. The storage device retains objective index information which associates primary key values with objective index values, and explanatory index information which associates values common to the primary keys with sets of values of explanatory indices of a plurality of items. The processor selects one or more items of the explanatory indices, clusters the values of the explanatory indices of the selected one or more items, identifies a range of the values of the explanatory indices of each of the items of each of the clusters which are obtained by the clustering, and outputs the identified value ranges.


