Deviation Analysis Feature Selection via Dissimilarity Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Performing deviation analysis on large datasets with many discrete features is computationally expensive and requires significant resources, making it inefficient.
Innovation Solution
An automated system that selects candidate discrete features based on dissimilarity scores, allowing deviation analysis to focus on features with significant deviational relationships, thereby reducing resource consumption and improving efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If deviation analysis is performed on all discrete features in large datasets, then comprehensive deviation relationships are identified, but computational resource consumption increases significantly
Solution Approach 1:
The patent segments the discrete features into two groups: candidate discrete features selected based on dissimilarity scores, and non-candidate discrete features excluded from analysis. This segmentation allows the system to focus computational resources only on the most promising features, thereby reducing overall resource consumption while maintaining analysis quality for the selected subset.
Solution Approach 2:
The patent performs preliminary selection of candidate discrete features using dissimilarity scores before conducting the actual deviation analysis. This preliminary action filters out features unlikely to show significant deviation relationships, so that the subsequent deviation analysis operates on a reduced set of features, reducing computational burden while preserving meaningful insights.
2Measurement precision
If deviation analysis is performed on all discrete features, then all potential deviation relationships are detected, but processing time increases
Solution Approach 1:
The patent segments the discrete features into candidate and non-candidate groups based on dissimilarity scores. By analyzing only the candidate subset, the system reduces processing time significantly while still detecting the most important deviation relationships, thus resolving the time-completeness contradiction.
Solution Approach 2:
The patent performs preliminary filtering of discrete features using dissimilarity scores before deviation analysis. This preliminary action eliminates features that are unlikely to exhibit meaningful deviation patterns, thereby reducing the processing time required for the main deviation analysis without missing critical insights.
3Productivity
If automated feature selection based on dissimilarity scores is implemented, then resource efficiency improves, but system complexity increases
Solution Approach 1:
The patent implements automated feature selection where the system itself determines which discrete features are most likely to show deviation relationships based on dissimilarity scores. This self-service approach eliminates the need for manual feature selection by analysts, improving resource efficiency and productivity, while the automation handles the complexity internally.
Solution Approach 2:
The patent changes the selection criterion from manual expert judgment to an automated parameter-based approach using dissimilarity scores. This parameter change enables systematic, reproducible feature selection that improves resource efficiency, while the computational calculation of dissimilarity scores manages the complexity in a standardized way.
Data Source
AI summary
Systems and methods include determination, for each of a plurality of discrete features, of statistics based on a number of occurrences of each discrete value of the discrete feature in the data, determination of first summary statistics based on the determined statistics, determine of a dissimilarity for each discrete feature based on the first summary statistics and on the statistics determined for the discrete feature, determination of candidate discrete features based on the determined dissimilarities, determination, for each of the candidate discrete features, of second summary statistics based on values of a continuous feature associated with each discrete value of the candidate discrete feature, determination of a deviation score for each of the candidate discrete features based on the second summary statistics, and transmission of the candidate discrete features for display in association with the continuous feature based on the determined deviation scores.


