Anomaly Cause Ranking in Large Data Sets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing volume of data generated by modern test equipment overwhelms test personnel, making it difficult for engineers to understand and analyze large datasets without advanced data processing systems, particularly in identifying measurement conditions that explain anomalies in measured values.
Innovation Solution
A method and system for real-time data processing that separates rows into target and remainder groups, computes relevance measures for measurement conditions, and displays significant associations, utilizing quantization to reduce computational workload and improve statistical accuracy, with a multi-processor system and user interface for efficient analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If engineers manually analyze large datasets generated by test equipment, then they can understand measurement conditions, but the analysis time and computational resources required become excessive
Solution Approach 1:
The patent segments the large dataset into manageable components by separating measurement conditions from measured values, and further segments analysis by dividing datasets into batches processed by multiple processors. This segmentation enables parallel processing while maintaining comprehensive anomaly detection capabilities.
Solution Approach 2:
The patent introduces an intermediary automated data processing system that acts as a mediator between the test equipment generating data and the engineers needing to analyze it. This intermediary system performs preliminary anomaly detection and prioritization, filtering out routine cases and presenting only significant anomalies to engineers for detailed analysis.
2Loss of information
If engineers manually identify measurement conditions explaining anomalies, then they can understand data causes, but the complexity of programming and data processing skills required increases significantly
Solution Approach 1:
The system enables self-service anomaly analysis by automatically computing relevance measures and generating prioritized lists of measurement conditions. The automated system serves itself by performing what would otherwise require skilled programmers to implement - calculating statistical relevance, processing large datasets, and presenting results in an engineer-friendly format without requiring engineers to write complex analysis code.
Solution Approach 2:
The patent transforms the complex multi-dimensional problem of anomaly analysis into a simplified parameter-based approach by computing a single relevance measure parameter for each measurement condition. This parameter transformation converts complex statistical relationships into a simple ranked list that engineers can easily interpret without needing advanced programming or statistical skills.
3Loss of information
If comprehensive analysis of all measurement conditions is performed, then all possible causes are identified, but the computational workload and processing time increase significantly
Solution Approach 1:
The patent applies partial action by computing relevance measures for only the most promising measurement conditions first, rather than exhaustively analyzing all possible conditions. The system processes measurement conditions in a prioritized sequence, stopping when sufficient causes are identified or when diminishing returns are reached, thus performing just enough analysis to solve the problem efficiently.
Solution Approach 2:
The patent introduces a relevance measure parameter that transforms the analysis from a brute-force enumeration of all measurement conditions into a targeted search based on statistical relevance. This parameter change enables the system to focus computational resources on the most likely causes, dramatically improving processing efficiency while maintaining comprehensive cause identification.
4Measurement precision
If statistical analysis is performed on large datasets to identify anomalies, then measurement conditions can be correlated with anomalies, but the computational resources and processing time required become prohibitive
Solution Approach 1:
The patent segments the computational workload by dividing the dataset into batches that can be processed in parallel by multiple processors. This segmentation reduces the memory footprint and computational burden on individual processors while maintaining the ability to perform comprehensive statistical analysis across the entire dataset through aggregation of partial results.
Solution Approach 2:
The patent transforms complex statistical correlation calculations into a more efficient parameter-based relevance measure that requires fewer computational resources. By changing from exhaustive correlation analysis to a streamlined relevance computation, the system achieves comparable analytical accuracy with significantly reduced energy consumption and processing time.
Data Source
AI summary
A method of operating a data processing system is disclosed. The method causes the data processing system to analyze a database that includes a measurement table characterized by one or more measurement condition columns and one or more measured parameter columns. Each row of the measurement table corresponding to an instance of the measured parameter values and the measurement condition values associated with the measured parameter values. The data processing system defines a target group that includes a plurality of rows in the measurement table and a remainder group that includes a plurality of rows in the measurement table that are different from the target group. The data processing system defines a measurement condition to be explored. A relevance measure is computed for each of the different measurement condition values and displayed.


