Feature Engineering via Interactive Structured-Unstructured Data Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In machine learning, feature engineering often faces challenges due to insufficient data from structured sources, leading to incomplete or unavailable features, which can impact model accuracy, especially in automated AI modeling pipelines.
Innovation Solution
A system and method that interactively learn features across structured and unstructured data sources using natural language processing (NLP) to augment the feature set by identifying and merging features from unstructured data sources, such as concept mapping and gap filling, to enhance model accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If features are extracted only from structured data sources, then the feature extraction process is simple and fast, but the feature completeness and model accuracy deteriorate due to insufficient data
Solution Approach 1:
The patent merges structured and unstructured data sources to extract features. The system combines tabular data from structured sources with text data from unstructured sources, processing both types together to generate comprehensive features that improve model accuracy while managing complexity through integrated processing pipelines.
Solution Approach 2:
The patent introduces NLP processing and feature mapping mechanisms as intermediaries between unstructured data and the machine learning model. These intermediaries transform unstructured text into structured features that can be combined with existing structured data, bridging the gap between different data types without requiring complex direct integration.
2Productivity
If automated AI modeling pipelines are used, then modeling efficiency is improved, but feature data availability deteriorates due to lack of manual intervention for data curation
Solution Approach 1:
The patent implements self-service mechanisms where the automated pipeline automatically identifies, extracts, and processes features from both structured and unstructured data sources without requiring manual intervention. The system autonomously handles data curation, feature engineering, and model training, maintaining high efficiency while preventing information loss through automated data discovery and processing capabilities.
3Quantity of substance
If unstructured data sources are incorporated for feature extraction, then feature richness is improved, but processing complexity and computational resources worsen
Solution Approach 1:
The patent extracts only the essential and relevant features from unstructured data sources using NLP techniques. Instead of processing all unstructured data in full detail, the system selectively extracts key information, concepts, and entities that are most valuable for the specific modeling task, reducing computational overhead while maintaining feature richness.
Solution Approach 2:
The patent applies partial processing to unstructured data by focusing computational resources on the most critical aspects of the data that contribute most to model performance. The system processes only the necessary portions of unstructured data required for feature extraction, avoiding excessive computation on redundant or less important information.
Data Source
AI summary
A concept associated with a feature used in machine learning model can be determined, the feature extracted from a first data source. A second data source containing the concept can be identified. An additional feature can be generated by performing a natural language processing on the second data source. The feature and the additional feature can be merged. A second machine learning model can be generated, which use the merged feature. A prediction result of the first machine learning model can be compared with a prediction result of the second machine learning model relative to ground truth data, to evaluate effective of the merged feature. Based on the evaluated effectiveness, the feature can be augmented with the merged feature in machine learning.


