Metadata Graph Feature Expansion for Automated Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current machine learning processes face challenges in optimizing model performance due to limited feature selection and noise reduction, which hinders their effectiveness in real-world applications.

Innovation Solution

The method involves leveraging metadata and graph techniques to identify and suggest supplemental features relevant to machine learning models, using a k-partite metadata graph to filter and expand features, thereby enhancing model performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional machine learning processes use limited feature selection, then the process remains simple and fast, but model performance is hindered due to insufficient feature information

Engineering Contradiction:
Improvemodel performanceVSAvoidfeature selection complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces a metadata graph as an intermediary structure that mediates between the raw feature data and the machine learning model. The metadata graph organizes feature information into a structured representation with nodes and edges, enabling complex feature relationships to be managed systematically without overwhelming the model processing capabilities

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments feature information into distinct metadata categories (e.g., data type, source, transformation rules) organized within the metadata graph structure. This segmentation allows the system to handle complex feature selection by breaking down the overall complexity into manageable, organized components that can be processed systematically

Inventive Principle:
Principle #1Segmentation

2Reliability

If traditional machine learning processes lack noise reduction mechanisms, then the processing remains simple, but effectiveness is reduced due to noise in feature data

Engineering Contradiction:
Improvepredictive capabilitiesVSAvoidnoise reduction complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by implementing noise reduction and feature validation within the metadata graph construction phase, before the data reaches the machine learning model. The metadata graph structure inherently organizes and filters noisy data through its structured relationships, performing cleanup operations in advance rather than during model training

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The metadata graph incorporates feedback mechanisms that continuously validate feature information against established metadata schemas and relationships. This feedback loop identifies and corrects noisy or inconsistent data entries, improving data quality through iterative validation rather than relying on complex post-processing noise reduction

Inventive Principle:
Principle #23Feedback

3Reliability

If the system expands features using comprehensive metadata graphs, then feature relevance improves, but processing time increases

Engineering Contradiction:
Improvefeature relevanceVSAvoidfeature processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements dynamics by making the metadata graph traversal adaptive to the specific machine learning task and data characteristics. The system dynamically adjusts the depth and breadth of metadata graph exploration based on task requirements, stopping expansion when sufficient relevant features are identified rather than exhaustively processing all possible metadata relationships

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies local quality by focusing metadata graph expansion on specific regions and relationships most relevant to the current machine learning task. Rather than uniformly processing the entire metadata graph, the system identifies and expands only the locally relevant portions that contribute most to feature relevance for the specific prediction problem

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12086144B2Metadata-based feature expansion for automated machine learning processes
Publication Date: 2024.09.10 DELL PROD LP
  • US12086144B2 patent drawing
  • US12086144B2 patent drawing
  • US12086144B2 patent drawing

AI summary

A method and system for metadata-based feature expansion for automated machine learning processes. In machine learning, feature selection—or a method through model inputs are reduced in dimensionality by retaining the relevant features, while also discarding the noise, thereof—often plays a pivotal role in the availability of algorithm(s) to derive model(s) from, the optimizing of said model(s) through training and/or testing, and, ultimately, the reaching of acceptable performance thresholds evaluating said model(s), which leads to their implementation in solving real-world problems. Embodiments disclosed herein, accordingly, leverage captured dataset metadata, as well as graph techniques, to identify and suggest one or more supplemental features distinct from original features for, and yet relevant to, the real-world problem and any machine learning models being evaluated.