Chained Feature Synthesis Dimensional Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The sheer number of combinations and confounding factors in real-world data can hinder pattern recognition and insight development in machine learning models, as uncontrolled feature synthesis often results in repetitive and noisy data without enhancing predictive capability.
Innovation Solution
Implementing feature extraction and dimension-reducing operations on datasets with synthesized features, using correlation matrices to identify and filter out redundant features, and generating a modified dataset based on a reduced feature set that satisfies a user-defined correlation threshold.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If deep feature synthesis is applied to generate new features from a dataset, then predictive capability and pattern recognition improve, but the dataset becomes filled with repetitive and noisy features that reduce data quality
Solution Approach 1:
The patent applies dimension-reducing operations that transform the feature space by changing parameters such as feature selection criteria and dimensionality metrics. This resolves the contradiction by modifying the dataset parameters to eliminate redundant features while preserving predictive capability.
Solution Approach 2:
The patent applies different processing treatments to different features based on their individual characteristics. By evaluating each feature's contribution and applying selective dimension reduction, it preserves high-quality predictive features while removing noisy repetitive ones, achieving local optimization of data quality.
2Loss of information
If automated feature generation subsystems synthesize new features using deep algorithms, then new insights and correlations are discovered, but the volume of data increases with corresponding noise without improving predictive capability
Solution Approach 1:
The patent extracts only the essential and informative features from the synthesized feature set by applying dimension-reducing operations. This separation process removes redundant and noisy features while retaining the valuable insights, thus reducing data volume without losing critical information.
Solution Approach 2:
The patent performs partial feature synthesis by applying dimension reduction to retain only the necessary portion of synthesized features. Instead of using all generated features, it selects a subset that provides sufficient predictive capability, avoiding the excess data volume problem.
3Productivity
If feature extraction and dimension-reducing operations are applied to reduce data dimensionality, then noise is eliminated and predictive capability improves, but the complexity of data processing operations increases
Solution Approach 1:
The patent applies dimension-reducing operations as a preliminary step before machine learning model training. By pre-processing the data to eliminate redundant features beforehand, it simplifies subsequent processing steps and improves overall productivity, despite the added initial processing complexity.
Data Source
AI summary
A method includes obtaining a first dataset including a first feature set, generating a first set of feature values by providing the first dataset to a set of feature primitive stacks, and determining a reduced set of feature values based on the first set of feature values by dimensionally reducing features of the first set of feature values. The method further includes generating an intermediate set of feature values by providing a value of the first dataset and a value of the reduced set of feature values to at least one feature primitive of the set of feature primitive stacks. The method further includes updating the reduced set of feature values by dimensionally reducing features of the intermediate set of feature values and storing a second dataset including features of the intermediate set of feature values in association with the first feature set.


