Feature Catalog Enhancement via Automated Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning applications face challenges in efficiently identifying relevant features of predictive models, as current feature catalogs often rely solely on metadata, requiring extensive research and analysis to distinguish between relevant and irrelevant features, leading to prolonged data analysis times.

Innovation Solution

A system and method that capture training data lineage metadata, model design time measurements, and runtime metrics to correlate with features in the feature catalog, expeditiously identifying relevant features by populating a feature catalog with features from predictive model training data and executing analyses to determine their impact on model predictions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If feature catalogs rely solely on metadata for identifying relevant features, then the system maintains simplicity in data storage, but the time and effort required to analyze and distinguish relevant features increases significantly

Engineering Contradiction:
Improvedata analysis timeVSAvoidfeature catalog structure
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The patent merges multiple data sources including training data lineage metadata, model design time measurements, model design time metadata, and runtime metrics into a unified feature catalog. This integration allows the system to automatically identify relevant features by correlating information across these different sources, thereby reducing the time and effort required for feature analysis without requiring manual intervention.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system performs preliminary actions by capturing and storing lineage metadata, design time measurements, and metadata during the model building process. This advance preparation ensures that when feature analysis is needed, the information is already organized and correlated in the feature catalog, eliminating the need for extensive retrospective research and analysis.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If extensive research and analysis are performed to distinguish relevant features from irrelevant features, then feature identification accuracy improves, but productivity decreases due to prolonged analysis times

Engineering Contradiction:
Improvefeature identification accuracyVSAvoidmodel building efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent implements feedback mechanisms by capturing runtime metrics and correlating them with feature catalog information. This feedback loop allows the system to continuously refine feature identification accuracy by learning from actual model performance data, while maintaining high productivity because the feedback is automatically collected and processed without requiring manual analysis.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs self-service by automatically correlating lineage metadata, design time measurements, and runtime metrics to identify relevant features. This automated self-analysis eliminates the need for manual research and analysis, thereby maintaining high feature identification accuracy while significantly improving model building efficiency and productivity.

Inventive Principle:
Principle #25Self-service

3Extent of automation

If manual analysis is used to determine relevant features, then interpretability is maintained, but automation level remains low requiring significant human involvement

Engineering Contradiction:
Improvefeature identification automationVSAvoidanalysis system complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent creates a universal system that handles multiple functions including capturing lineage metadata, storing design time measurements, collecting runtime metrics, and correlating all this information to automatically identify relevant features. This multi-functional automated system replaces manual analysis processes while maintaining interpretability through the structured correlation of data sources, thereby significantly increasing the extent of automation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11651281B2Feature catalog enhancement through automated feature correlation
Publication Date: 2023.05.16 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11651281B2 patent drawing
  • US11651281B2 patent drawing
  • US11651281B2 patent drawing

AI summary

Embodiments relate to a system, program product, and method for generating an enhanced feature catalog for a predictive model. The embodiments disclosed herein include capturing predictive model design time information including training data lineage metadata to determine the features of the training data, model design time measurements, and model design time metadata. Once the predictive model is built, the training data lineage metadata is used to capture the features that will be maintained within a feature catalog. The model design time measurements and model design time metadata provide further correlation between the predictive model and the features. Runtime metrics on the predictive model create additional correlations between the captured data and metadata with the features in the feature catalog to expeditiously identify the relevant features of the predictive model.