Sensor Data Imputation via Correlated Response Sequences
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for imputing missing sensor data rely heavily on previous history, which is often difficult to obtain, leading to information loss and increased imputation errors, particularly in critical applications like healthcare and smart infrastructure.
Innovation Solution
A method that identifies a similar response data sequence with high correlation to the sensor data sequence, determines nearest neighbors within a nearness threshold, and uses semantics-based learning to estimate missing data values without relying on previous sensor data history, employing mechanisms like Missing Completely at Random (MCAR) and Missing at Random (MAR).
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional methods estimate missing data based on previous history, then imputation can be performed, but previous history is difficult to obtain and information loss occurs
Solution Approach 1:
The patent introduces response data sequences as intermediary elements that mediate between available sensor data and missing historical data. Instead of directly accessing unavailable historical information, the system uses response data sequences that are correlated with the sensor data to indirectly infer missing values, thus avoiding information loss while maintaining imputation accuracy
Solution Approach 2:
The patent replaces the conventional mechanical approach of directly accessing historical data with a computational substitution method. It uses correlation-based identification of response data sequences and semantic learning algorithms to substitute for the unavailable historical data access mechanism, achieving imputation through data relationships rather than direct historical retrieval
2Ease of operation
If conventional methods use previous sensor data history for imputation, then missing values can be estimated, but imputation errors increase when historical data is unavailable
Solution Approach 1:
The patent changes the parameters used for imputation from direct historical data values to correlation-based response data sequence characteristics. By identifying response data sequences with high correlation coefficients and using their semantic features, the system maintains imputation feasibility without relying on unavailable historical data, thereby reducing imputation errors
Solution Approach 2:
The patent creates a computational copy of the historical data relationship through response data sequences. Instead of accessing the original historical data directly, it identifies response sequences that replicate the informational content and relationships of historical data, allowing imputation to proceed with the same feasibility and precision as if historical data were available
3Measurement precision
If the method uses semantics-based learning to estimate missing data, then imputation accuracy improves, but the complexity of data processing increases
Solution Approach 1:
The patent segments the complex semantics-based learning process into distinct modular steps: (1) identification of response data sequences based on correlation thresholds, (2) extraction of semantic features from identified sequences, (3) application of learning algorithms to estimate missing values. This segmentation reduces processing complexity by making each step independent and manageable while maintaining overall imputation accuracy
Data Source
AI summary
Embodiments herein provide a method for imputing sensor data, in a sensor data sequence with missing data based on the semantics learning, where semantics is defined by the constraints of the sensor data features. A candidate value for imputation is determined based on sensor data of corresponding instances of time instants of the sensor data sequence using learning based on semantics of features of the sensor data sequence with missing data. The nearest neighbors search has been applied in similar response data sequence using the data values corresponding to the time instant of missing data in sensor data sequence. In case similar response data sequence is not available imputation is performed based on the distribution pattern of missing data.


