Data Analyzing Device Using Supplementary Information to Narrow Feature Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data mining techniques generate a large volume of features through automatic generation, making it time-consuming to design optimal features and select relevant ones, especially when dealing with noisy data, as they produce features with high correlations due to the use of multiple arithmetic operators without considering the meaning of each data column.
Innovation Solution
A data analyzing device and method that includes a data input unit, display unit, supplementary information adding unit, rule storage unit, and prediction model generating unit, where supplementary information is added to features based on user input to determine whether to apply calculation operations for generating new features, using fixed and additional rules to narrow down feature generation and selection efficiently.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If automatic generation of features using multiple arithmetic operators is performed, then the quantity of generated features increases, but the time required for feature generation and selection increases
Solution Approach 1:
The patent applies local quality by differentiating the treatment of features based on their characteristics. Features are categorized into those with physical meaning and those without, and different calculation operations are applied selectively. This localized differentiation reduces unnecessary feature generation while maintaining relevant features, thereby reducing overall processing time.
Solution Approach 2:
The patent changes the parameter of feature generation by introducing a filter based on physical meaning. Instead of generating all possible combinations, the system modifies the generation process to exclude features that lack physical meaning, thus reducing the quantity of features to be processed and decreasing time requirements.
2Quantity of substance
If automatic generation of features using multiple arithmetic operators is performed, then the quantity of generated features increases, but the difficulty of understanding the generated features increases
Solution Approach 1:
The patent extracts and removes features that lack physical meaning from the generated feature set. By taking out irrelevant features, the system leaves only features that are both numerous and meaningful, thereby reducing the difficulty of understanding while maintaining the quantity of useful features.
Solution Approach 2:
The patent applies local quality by differentiating the treatment of features based on their characteristics. Features are categorized into those with physical meaning and those without, and different calculation operations are applied selectively. This localized differentiation reduces unnecessary feature generation while maintaining relevant features, thereby reducing overall processing time.
3Quantity of substance
If feature selection is performed to narrow down features, then the number of features is reduced, but the time required for feature selection increases
Solution Approach 1:
The patent applies preliminary action by pre-filtering features based on physical meaning before the main feature selection process. This preliminary sorting reduces the number of candidates that need to be evaluated in detail, thereby reducing the time required for feature selection while maintaining an appropriate number of features.
Solution Approach 2:
The patent changes the parameter of feature generation by introducing a filter based on physical meaning. Instead of generating all possible combinations, the system modifies the generation process to exclude features that lack physical meaning, thus reducing the quantity of features to be processed and decreasing time requirements.
Data Source
AI summary
To enable effectively narrowing down features to be generated, thereby generating effective features at a high speed, in obtaining the features from a large volume of data. A fixed rule and an additional rule are stored in advance. The fixed rule specifies a rule of a calculation operation for generating a new feature. The additional rule specifies whether to perform a calculation operation for generating the new feature on a basis of meta-information, not depending on whether the fixed rule is applicable. An objective variable is predicted from plurality of features on the basis of the fixed rule and the additional rule.


