Automated Feature Engineering via Semantic Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current automated machine learning algorithms face difficulties in identifying relevant features, comparing them, and creating new features through feature engineering, leading to high costs and time consumption in preparing datasets for training machine learning systems, especially with large datasets.

Innovation Solution

A method and system that automatically perform feature engineering by determining similarity between features, clustering similar features, generating new features, and adding them to the dataset, which includes using metadata, semantic, and unit of measurement similarity values to improve dataset quality for machine learning systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual feature engineering is performed, then feature quality and relevance can be improved, but time consumption and labor costs increase significantly

Engineering Contradiction:
Improvefeature qualityVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs feature engineering automatically by analyzing dataset metadata, generating similarity values between features, clustering similar features, and creating new features through automated functions - eliminating the need for manual intervention while maintaining feature quality

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical feature engineering processes with an automated computational system that uses algorithms to calculate similarity values, cluster features, and generate new features, substituting human expertise with automated mechanical processes

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If more features are added to the dataset, then predictive accuracy may improve, but computational costs and processing time increase

Engineering Contradiction:
Improvepredictive accuracyVSAvoidcomputational cost
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system merges similar features into clusters based on calculated similarity values, creating consolidated feature groups that reduce redundancy while maintaining predictive information, thereby reducing the total number of features that need to be processed

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The automated system extracts and identifies only the most relevant features through similarity analysis and clustering, filtering out redundant or less informative features before they are used in training, thus reducing computational burden while maintaining accuracy

Inventive Principle:
Principle #2Taking out (Extraction)

3Adaptability or versatility

If features are manually selected and engineered, then feature relevance can be optimized, but labor resources and costs increase

Engineering Contradiction:
Improvefeature relevanceVSAvoidlabor requirements
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system autonomously analyzes dataset metadata, calculates similarity values between features, identifies relevant features through clustering, and generates new features automatically - eliminating the need for human labor while optimizing feature relevance through computational methods

Inventive Principle:
Principle #25Self-service

Data Source

PatentEP4372631A1Automated custom feature engineering
Publication Date: 2024.05.22 FUJITSU LTD
  • EP4372631A1 patent drawingFigure 1
  • EP4372631A1 patent drawingFigure 2
  • EP4372631A1 patent drawingFigure 3A

AI summary

A method may include obtaining a dataset that may include one or more columns, wherein each of the one or more columns may include a title and at least one value. The operations may further include extracting, for each of the one or more columns, the title and a sample value from the at least one value. The operations may additionally include, synthesizing a question based on the title and the sample value for each of the one or more columns. Further, the operations may include sending the question to a language model to obtain an answer. The operations may additionally include generating from the answer to the question, a predicted unit of measurement for the at least one value in each of the one or more columns. Systems and devices for performing the method are also disclosed.