Machine Learning Feature Nullification for Overlooked Pattern Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current machine learning approaches for analyzing large datasets tend to overlook statistically insignificant but operationally important features, which are crucial for investigations, as they focus on statistically significant variables, thereby missing important insights.
Innovation Solution
The proposed method, 'small factor analysis,' involves deriving an initial model using machine learning and then nullifying the contribution of statistically significant features to highlight overshadowed, less significant features, allowing for a re-optimized alternate model that brings these features to light, and presenting them to the user through a user interface for better investigation insights.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If machine learning focuses on statistically significant variables, then prediction accuracy is improved, but operationally important but statistically insignificant features are overlooked
Solution Approach 1:
The patent segments the feature analysis into two distinct phases: first analyzing statistically significant features for prediction accuracy, then separately analyzing statistically insignificant features for operational importance. This segmentation allows both types of features to be examined without one overwhelming the other, resolving the contradiction by giving each feature type dedicated analytical attention.
Solution Approach 2:
The patent applies inversion by nullifying the contribution of statistically significant features and re-running the machine learning process to highlight statistically insignificant features. This reverse approach allows operationally important but statistically insignificant features to emerge that would otherwise be overshadowed by dominant significant features.
2Loss of information
If manual analysis is used to identify operationally important features, then comprehensive insights are achieved, but time consumption increases significantly
Solution Approach 1:
The patent introduces an intermediary computational process that automatically identifies and highlights statistically insignificant features by nullifying significant features and re-running analysis. This intermediary machine learning process bridges the gap between automated analysis speed and comprehensive manual inspection depth, providing thorough feature identification without manual time consumption.
3Loss of information
If all features are analyzed equally, then no features are overlooked, but the analysis becomes computationally inefficient and difficult to interpret
Solution Approach 1:
The patent segments feature analysis by separating statistically significant features from statistically insignificant features through nullification and re-analysis. This segmentation creates distinct analytical pathways that reduce overall complexity while ensuring comprehensive feature coverage, as each segment can be analyzed with appropriate methods rather than applying uniform complex analysis to all features.
Solution Approach 2:
By inverting the standard approach of analyzing all features together, the patent nullifies significant features to isolate and analyze insignificant ones. This inversion simplifies the analysis of each feature subset by removing confounding influences from other features, making the analysis more interpretable while maintaining comprehensive coverage.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method, and corresponding system and computer program product, is provided for identifying meaningful information in connection with an investigation. The method comprises processing a dataset using a machine learning process to derive an initial model conveying first statistical significance information corresponding to features in the dataset. The method also comprises deriving an alternate model at least in part by processing the dataset using the machine learning process while nullifying a contribution of certain features in the dataset selected as candidates for nullification. The alternate model conveys second statistical significance information corresponding to features in the dataset. A user interface is rendered on the display and presents information for assisting the user in identifying the information in the dataset meaningful to the investigation, the information presented being derived at least in part by processing information conveyed by the initial model and the alternate model.