Weirdness Score Calculation for Data Anomaly Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data analysis systems require significant computational resources and are complex, limiting their ability to efficiently analyze large volumes of data and compare different types of data, often resulting in false positives due to failure to account for overall patterns or trends.
Innovation Solution
A computer-based method and system that calculates a 'weirdness score' for variables by comparing measured values to predicted values over time, ranking parameters relative to peers, and weighting ranks to identify anomalous data without requiring complex modeling or significant computational resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If complex modeling algorithms are used to analyze data, then analysis accuracy is improved, but computational resources and implementation complexity increase significantly
Solution Approach 1:
The patent transforms the analysis approach by changing from complex modeling algorithms to a parameter-based deviation measurement system. Instead of using sophisticated models, the system calculates predicted values based on historical data patterns and measures deviations from these predictions, achieving accurate anomaly detection with simpler computational methods
Solution Approach 2:
The patent replaces complex algorithmic modeling with a more straightforward mathematical approach involving predicted value calculation and deviation measurement. This substitution eliminates the need for complex modeling algorithms while maintaining analytical effectiveness through simpler computational operations
2Measurement precision
If complex modeling algorithms are used to analyze data, then analysis accuracy is improved, but processing time increases
Solution Approach 1:
The patent changes the computational approach from complex modeling to parameter-based deviation calculation. By using predicted values derived from historical data patterns and measuring deviations from these predictions, the system achieves accurate analysis with significantly reduced processing time compared to complex modeling algorithms
3Device complexity
If identical types of data are compared, then analysis simplicity is maintained, but false positives increase due to failure to account for overall patterns
Solution Approach 1:
The patent creates a universal analysis framework that can handle different types of data (transactional, operational, behavioral) using the same deviation-based approach. This multi-functional system compares data points against predicted values derived from historical patterns, enabling accurate anomaly detection across diverse data types without increasing complexity
Solution Approach 2:
The patent implements feedback mechanisms by continuously comparing actual data values against predicted values derived from historical patterns. This feedback loop allows the system to adapt to changing patterns and distinguish between normal variations and true anomalies, reducing false positives while maintaining analytical simplicity
4Productivity
If data parameters are analyzed in isolation, then processing efficiency is maintained, but anomaly detection accuracy decreases due to missed contextual patterns
Solution Approach 1:
The patent merges the analysis of multiple data parameters by calculating predicted values that incorporate historical patterns across different parameters. Instead of analyzing parameters in isolation, the system integrates contextual information from historical data to create comprehensive predictions, improving anomaly detection accuracy while maintaining processing efficiency through unified calculation approaches
Data Source
AI summary
A computer-based method of determining a weirdness score for variables within a data set is provided. The method includes receiving a selection of a first variable, wherein the first variable is defined by a measure, a time period, and a plurality of entities, calculating a plurality of parameters for the first variable, wherein each of the plurality of parameters is calculated based at least in part on a deviation of a measured value from a predicted value, calculating a rank for each of the plurality of parameters for the first variable, wherein the rank of each parameter is calculated relative to corresponding parameters calculated for all other variables in the data set having the same measure and time period as the first variable, and calculating a weirdness score for the first variable based at least in part on the calculated rank of each of the plurality of parameters.


