Ontology-Based Feature Engineering via Reinforcement Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing feature selection systems for machine learning algorithms are inefficient and impractical for non-experts, as they require manual intervention and are costly to validate, especially when dealing with large datasets and complex feature transformations.
Innovation Solution
An autonomous deep reinforcement learning network that iteratively generates and evaluates features based on statistical significance and interpretability, using a domain ontology to ensure that the features are both statistically important and interpretable by domain experts, thereby automating the feature engineering process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If manual feature selection is performed by data scientists using intuition and domain knowledge, then feature interpretability is improved, but productivity is worsened due to the time-consuming trial and error process
Solution Approach 1:
The system performs self-service by automatically generating features through transformations of existing features without requiring manual data scientist intervention. The automated feature generation engine continuously creates new features by applying transformations to existing features, eliminating the need for manual trial and error while maintaining feature quality through algorithmic selection criteria.
Solution Approach 2:
The patent replaces the mechanical manual process of feature selection with an automated computational system. Instead of data scientists manually creating and testing features, the system uses an automated feature generation engine that applies transformations and uses selection criteria to identify valuable features, substituting human manual work with algorithmic automation.
2Measurement precision
If exhaustive feature validation is performed by training and evaluating models for each newly-constructed feature, then measurement precision is improved, but loss of time is worsened due to the computational cost
Solution Approach 1:
The system applies partial validation by using selection criteria that evaluate features without requiring complete model training and evaluation for each feature. The automated feature generation engine uses efficient selection criteria to assess feature quality, performing only the necessary validation to determine feature value rather than exhaustive testing, thus reducing validation time while maintaining adequate feature quality assessment.
3Adaptability or versatility
If the number of possible features is increased to capture all potential patterns in data, then adaptability is improved, but device complexity is worsened due to the unlimited feature space
Solution Approach 1:
The system segments the feature generation process into manageable components: existing features serve as base elements, transformations are applied in systematic ways to generate new features, and selection criteria filter the results. This segmentation of the feature space into structured transformations and selection stages makes the complexity manageable while still exploring a comprehensive feature space.
Solution Approach 2:
The feature generation process is dynamic and iterative rather than static. The automated feature generation engine continuously generates new features through transformations and applies selection criteria to identify valuable features. This dynamic approach allows the system to adaptively explore the feature space over time, managing complexity through iterative refinement rather than attempting to handle all features simultaneously.
Data Source
AI summary
Systems and methods include generation of a first plurality of features using a learning network, determination of an interpretability of each of the first plurality of features based on a domain ontology and on symbolic rules associated with entities of the domain ontology, determination of a first set of the first plurality of features which were determined as interpretable, determination of a performance of a model trained using the first set of the plurality of features, determine a reward based on the performance and the interpretability, and generation of a second plurality of features using the learning network based on the reward.


