Automatic Feature Extraction from Relational Databases

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Feature engineering in predictive modeling is a bottleneck, consuming up to 80% of the effort in predictive analytics projects, as it requires manual domain knowledge to create relevant features for machine learning algorithms from relational databases.

Innovation Solution

A system that automatically extracts features from relational databases by generating an entity graph, joining tables based on relationships, and using data mining algorithms selected based on the type of data, facilitating faster and more efficient feature engineering.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual feature engineering is performed using domain knowledge, then feature quality and model accuracy improve, but time consumption and effort increase significantly (up to 80% of project effort)

Engineering Contradiction:
Improvefeature qualityVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables self-service feature engineering by automatically extracting features from relational databases using data mining algorithms. The automated feature extraction process eliminates the need for manual domain knowledge intervention, reducing time consumption while maintaining feature quality through algorithmic pattern recognition and statistical analysis.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the manual mechanical process of feature engineering with an automated computational system. Data mining algorithms substitute human analysts, automatically discovering patterns and extracting features from database tables, thereby reducing effort from 80% to minimal supervision while preserving feature quality through rigorous statistical methods.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated feature extraction is implemented, then productivity and processing efficiency improve, but the complexity of the system increases

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the feature engineering process into distinct automated components: data mining algorithm selection, feature extraction modules, and model training integration. This segmentation allows each component to be independently optimized and managed, reducing overall system complexity while maintaining high productivity through specialized automated processing for each stage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a universal automated feature extraction platform that handles multiple data types and mining algorithms through a single integrated system. This multi-functional approach consolidates what would otherwise require separate manual processes, improving productivity while managing complexity through unified architecture rather than multiple specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of information

If comprehensive data mining algorithms are applied to all data types, then feature extraction completeness improves, but computational resources and processing time increase

Engineering Contradiction:
Improvefeature completenessVSAvoidcomputational resources
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The system applies local quality by selecting and applying specific data mining algorithms tailored to each data type and column characteristics rather than uniformly applying all algorithms. This targeted approach ensures complete feature extraction for each specific data context while optimizing computational resource usage by avoiding unnecessary processing with inappropriate algorithms.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent dynamically adjusts processing parameters based on data type identification, selecting appropriate mining algorithms and extraction depth for each column. This parameter adaptation ensures comprehensive feature extraction where needed while reducing computational overhead for data types requiring less intensive processing, thereby balancing completeness with resource efficiency.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11645311B2Automatic feature extraction from a relational database
Publication Date: 2023.05.09 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11645311B2 patent drawing
  • US11645311B2 patent drawing
  • US11645311B2 patent drawing

AI summary

Techniques facilitating automatic feature extraction from a relational database are provided. In an embodiment, a method can include generating an entity graph based on a relational database, wherein the entity graph comprises a first node associated with a first table in the relational database and a second node associated with a second table in the relational database. In another embodiment, the method can include joining the first table and the second table based on an edge between the first table and the second table defined by the entity graph, wherein a resulting joined table is connected by a column of data. In another embodiment, the method can include extracting a feature from the column of data using a data mining algorithm selected from a set of data mining algorithms based on a type of data in the column of data.