Feature Catalog Enhancement via Automated Correlation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning applications face challenges in efficiently identifying relevant features of predictive models, as current feature catalogs often rely solely on metadata, requiring extensive research and analysis to distinguish between relevant and irrelevant features, leading to prolonged data analysis times.
Innovation Solution
A system and method that capture training data lineage metadata, model design time measurements, and runtime metrics to correlate with features in the feature catalog, expeditiously identifying relevant features by populating a feature catalog with features from predictive model training data and executing analyses to determine their impact on model predictions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If feature catalogs rely solely on metadata for identifying relevant features, then the system maintains simplicity in data storage, but the time and effort required to analyze and distinguish relevant features increases significantly
Solution Approach 1:
The patent merges multiple data sources including training data lineage metadata, model design time measurements, model design time metadata, and runtime metrics into a unified feature catalog. This integration allows the system to automatically identify relevant features by correlating information across these different sources, thereby reducing the time and effort required for feature analysis without requiring manual intervention.
Solution Approach 2:
The system performs preliminary actions by capturing and storing lineage metadata, design time measurements, and metadata during the model building process. This advance preparation ensures that when feature analysis is needed, the information is already organized and correlated in the feature catalog, eliminating the need for extensive retrospective research and analysis.
2Measurement precision
If extensive research and analysis are performed to distinguish relevant features from irrelevant features, then feature identification accuracy improves, but productivity decreases due to prolonged analysis times
Solution Approach 1:
The patent implements feedback mechanisms by capturing runtime metrics and correlating them with feature catalog information. This feedback loop allows the system to continuously refine feature identification accuracy by learning from actual model performance data, while maintaining high productivity because the feedback is automatically collected and processed without requiring manual analysis.
Solution Approach 2:
The system performs self-service by automatically correlating lineage metadata, design time measurements, and runtime metrics to identify relevant features. This automated self-analysis eliminates the need for manual research and analysis, thereby maintaining high feature identification accuracy while significantly improving model building efficiency and productivity.
3Extent of automation
If manual analysis is used to determine relevant features, then interpretability is maintained, but automation level remains low requiring significant human involvement
Solution Approach 1:
The patent creates a universal system that handles multiple functions including capturing lineage metadata, storing design time measurements, collecting runtime metrics, and correlating all this information to automatically identify relevant features. This multi-functional automated system replaces manual analysis processes while maintaining interpretability through the structured correlation of data sources, thereby significantly increasing the extent of automation.
Data Source
AI summary
Embodiments relate to a system, program product, and method for generating an enhanced feature catalog for a predictive model. The embodiments disclosed herein include capturing predictive model design time information including training data lineage metadata to determine the features of the training data, model design time measurements, and model design time metadata. Once the predictive model is built, the training data lineage metadata is used to capture the features that will be maintained within a feature catalog. The model design time measurements and model design time metadata provide further correlation between the predictive model and the features. Runtime metrics on the predictive model create additional correlations between the captured data and metadata with the features in the feature catalog to expeditiously identify the relevant features of the predictive model.


