Automated Data Model Updating for Scalable Sensor Fault Cleansing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data cleansing systems for large-scale industrial processes, such as oil production facilities, face scalability issues and are inadequate in detecting and addressing errors in data from isolated sensors without significant correlations with other instruments, leading to inaccurate process monitoring and optimization.
Innovation Solution
An automated method and system for building and maintaining models that perform clustering on data sources, building data models, and automatically cleansing operational data using dynamic principal components analysis (DPCA) and recursive methods, allowing for the detection and correction of faulty data across multiple and isolated sensors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data driven models including PCA or PLS are used to monitor process statistics to detect sensor failures, then fault detection capability is improved, but scalability to large-scale data collection systems deteriorates
Solution Approach 1:
The patent segments the large-scale data collection system into multiple clusters, each monitored by its own data driven model (PCA or PLS). This segmentation allows the system to maintain fault detection capability for each cluster while reducing the overall complexity and improving scalability, as each model only needs to process a subset of the total data rather than all data simultaneously.
2Measurement precision
If correlation between input variables is used for fault detection, then detection accuracy for correlated sensors is improved, but applicability to isolated sensors without correlations deteriorates
Solution Approach 1:
The patent creates a universal fault detection framework that can handle both correlated and isolated sensors. By organizing sensors into clusters where correlation-based models are applied within clusters, and isolated sensors are monitored individually, the system achieves universal applicability across different sensor types and configurations while maintaining detection accuracy for each case.
Data Source
AI summary
Methods and systems for building and maintaining model(s) of a physical process are disclosed. One method includes receiving training data associated with a plurality of different data sources, and performing a clustering process to form one or more clusters. For each of the one or more clusters, the method includes building a data model based on the training data associated with the data sources in the cluster, automatically performing a data cleansing process on operational data based on the data model, and automatically updating the data model based on updated training data that is received as operational data. For data sources excluded from the clusters, automatic building, data cleansing, and updating of models can also be applied.


