Automated Feature Engineering via Semantic Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current automated machine learning algorithms face difficulties in identifying relevant features, comparing them, and creating new features through feature engineering, leading to high costs and time consumption in preparing datasets for training machine learning systems, especially with large datasets.
Innovation Solution
A method and system that automatically perform feature engineering by determining similarity between features, clustering similar features, generating new features, and adding them to the dataset, which includes using metadata, semantic, and unit of measurement similarity values to improve dataset quality for machine learning systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual feature engineering is performed, then feature quality and relevance can be improved, but time consumption and labor costs increase significantly
Solution Approach 1:
The system performs feature engineering automatically by analyzing dataset metadata, generating similarity values between features, clustering similar features, and creating new features through automated functions - eliminating the need for manual intervention while maintaining feature quality
Solution Approach 2:
The patent replaces manual mechanical feature engineering processes with an automated computational system that uses algorithms to calculate similarity values, cluster features, and generate new features, substituting human expertise with automated mechanical processes
2Reliability
If more features are added to the dataset, then predictive accuracy may improve, but computational costs and processing time increase
Solution Approach 1:
The system merges similar features into clusters based on calculated similarity values, creating consolidated feature groups that reduce redundancy while maintaining predictive information, thereby reducing the total number of features that need to be processed
Solution Approach 2:
The automated system extracts and identifies only the most relevant features through similarity analysis and clustering, filtering out redundant or less informative features before they are used in training, thus reducing computational burden while maintaining accuracy
3Adaptability or versatility
If features are manually selected and engineered, then feature relevance can be optimized, but labor resources and costs increase
Solution Approach 1:
The system autonomously analyzes dataset metadata, calculates similarity values between features, identifies relevant features through clustering, and generates new features automatically - eliminating the need for human labor while optimizing feature relevance through computational methods
Data Source
Figure 1
Figure 2
Figure 3A
AI summary
A method may include obtaining a dataset that may include one or more columns, wherein each of the one or more columns may include a title and at least one value. The operations may further include extracting, for each of the one or more columns, the title and a sample value from the at least one value. The operations may additionally include, synthesizing a question based on the title and the sample value for each of the one or more columns. Further, the operations may include sending the question to a language model to obtain an answer. The operations may additionally include generating from the answer to the question, a predicted unit of measurement for the at least one value in each of the one or more columns. Systems and devices for performing the method are also disclosed.