Feature Engineering via Interactive Structured-Unstructured Data Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In machine learning, feature engineering often faces challenges due to insufficient data from structured sources, leading to incomplete or unavailable features, which can impact model accuracy, especially in automated AI modeling pipelines.

Innovation Solution

A system and method that interactively learn features across structured and unstructured data sources using natural language processing (NLP) to augment the feature set by identifying and merging features from unstructured data sources, such as concept mapping and gap filling, to enhance model accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If features are extracted only from structured data sources, then the feature extraction process is simple and fast, but the feature completeness and model accuracy deteriorate due to insufficient data

Engineering Contradiction:
Improvefeature completenessVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges structured and unstructured data sources to extract features. The system combines tabular data from structured sources with text data from unstructured sources, processing both types together to generate comprehensive features that improve model accuracy while managing complexity through integrated processing pipelines.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces NLP processing and feature mapping mechanisms as intermediaries between unstructured data and the machine learning model. These intermediaries transform unstructured text into structured features that can be combined with existing structured data, bridging the gap between different data types without requiring complex direct integration.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If automated AI modeling pipelines are used, then modeling efficiency is improved, but feature data availability deteriorates due to lack of manual intervention for data curation

Engineering Contradiction:
Improvemodeling efficiencyVSAvoidfeature data availability
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent implements self-service mechanisms where the automated pipeline automatically identifies, extracts, and processes features from both structured and unstructured data sources without requiring manual intervention. The system autonomously handles data curation, feature engineering, and model training, maintaining high efficiency while preventing information loss through automated data discovery and processing capabilities.

Inventive Principle:
Principle #25Self-service

3Quantity of substance

If unstructured data sources are incorporated for feature extraction, then feature richness is improved, but processing complexity and computational resources worsen

Engineering Contradiction:
Improvefeature richnessVSAvoidcomputational resources
Core Design Contradiction:
Quantity of substanceVSUse of energy by moving object

Solution Approach 1:

The patent extracts only the essential and relevant features from unstructured data sources using NLP techniques. Instead of processing all unstructured data in full detail, the system selectively extracts key information, concepts, and entities that are most valuable for the specific modeling task, reducing computational overhead while maintaining feature richness.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial processing to unstructured data by focusing computational resources on the most critical aspects of the data that contribute most to model performance. The system processes only the necessary portions of unstructured data required for feature extraction, avoiding excessive computation on redundant or less important information.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20230029218A1Feature engineering using interactive learning between structured and unstructured data
Publication Date: 2023.01.26 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20230029218A1 patent drawing
  • US20230029218A1 patent drawing
  • US20230029218A1 patent drawing

AI summary

A concept associated with a feature used in machine learning model can be determined, the feature extracted from a first data source. A second data source containing the concept can be identified. An additional feature can be generated by performing a natural language processing on the second data source. The feature and the additional feature can be merged. A second machine learning model can be generated, which use the merged feature. A prediction result of the first machine learning model can be compared with a prediction result of the second machine learning model relative to ground truth data, to evaluate effective of the merged feature. Based on the evaluated effectiveness, the feature can be augmented with the merged feature in machine learning.