Anomaly Detection in High-Dimensional Streaming Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manufacturing analytics face challenges in diagnosing partial equipment failure or degradation due to the complexity and magnitude of streaming data from millions of parts and assembly lines, exacerbated by the lack of labeled data, which limits the effectiveness of supervised learning algorithms.
Innovation Solution
A method for detecting anomalous data using a trained model that applies Hotelling's T2 statistics and Q-residual to clean up outliers, calculating principal components, and deploying machine learning or statistical models to edge and cloud infrastructure for real-time anomaly detection in manufacturing lines, leveraging statistical analysis and online inferential sensing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If supervised learning algorithms are applied to manufacturing data, then detection accuracy can be improved, but the lack of labeled data makes it impossible to implement
Solution Approach 1:
The patent introduces unsupervised learning algorithms as an intermediary approach between traditional rule-based methods and supervised learning. These algorithms process unlabeled manufacturing data to automatically identify anomalies without requiring manual labeling, thereby achieving detection accuracy comparable to supervised methods while bypassing the labeled data bottleneck
Solution Approach 2:
The system enables manufacturing data to serve itself by using unsupervised learning algorithms that automatically discover patterns and anomalies in the data without external labeling or intervention. The algorithms self-train on the available unlabeled data, performing anomaly detection autonomously without requiring human-annotated datasets
2Ease of manufacture
If traditional statistical methods are used for anomaly detection, then implementation is simple, but they cannot handle high-dimensional streaming data effectively
Solution Approach 1:
The patent transforms the approach to handling high-dimensional data by using unsupervised learning algorithms that can operate in high-dimensional spaces without requiring dimensionality reduction or simplification. These algorithms naturally accommodate the complexity of streaming manufacturing data across multiple sensors and parameters, maintaining implementation feasibility while handling increased dimensionality
Solution Approach 2:
The system changes the fundamental parameters of anomaly detection by transitioning from traditional statistical thresholds to unsupervised learning-based anomaly scoring. This parameter transformation allows the system to handle high-dimensional streaming data effectively while maintaining relative implementation simplicity through automated model training and deployment
3Reliability
If more sensors and data collection points are added to manufacturing lines, then monitoring coverage is improved, but data complexity and processing difficulty increase
Solution Approach 1:
The patent implements a universal unsupervised learning framework that can process data from multiple sensors and data sources simultaneously. This multi-functional approach allows the same anomaly detection system to handle diverse data types and dimensions from various manufacturing line components, improving monitoring coverage without proportionally increasing processing complexity
Solution Approach 2:
The system segments the complex high-dimensional data processing task into manageable components through unsupervised learning algorithms that automatically identify and process different data patterns independently. This segmentation allows the system to handle data from numerous sensors effectively by breaking down the overall complexity into autonomous processing streams
Data Source
AI summary
Method for detecting anomalous data in a manufacturing line or live sensing application. The method includes computing a projection of new incoming data on a trained model and identifying potential anomalies by comparing a window for the incoming data to normal representation criteria based upon user-specified thresholds. The trained model is created by applying hoteling T2 statistics and Q-residual to clean up outliers from an historic time interval of data and calculating principal components of the data and choosing a subset of components which represent a variability in the data. A model deployment pipeline is generated from the trained model and which is capable of deploying machine learning or statistical models to an edge and cloud infrastructure associated with the manufacturing line or live sensing application.


