Machine Learning Audit Trail for Reversible Sensor Time-Series
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems fail to ensure auditability of machine learning model results in sensor-based automation, leading to issues with data integrity and compliance in regulated industries, where large volumes of time-series data from sensors are prone to anomalies and tampering, making it difficult to recover original source data for legal and regulatory purposes.
Innovation Solution
The implementation of a system that creates a tamper-proof data set through intelligent preprocessing and Multivariate State Estimation Technique (MSET), which ensures determinism, compression, and reversibility, allowing for the recovery of original data and detection of any tampering, while reducing adversarial relationships between industries and regulators.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If sensor-based automation systems store large volumes of time-series data for machine learning analysis, then the system can enable prognostics and anomaly detection, but the data becomes vulnerable to corruption and tampering, making auditability difficult
Solution Approach 1:
The system performs preliminary actions by creating compressed representations of sensor data and recording provenance metadata before data corruption or tampering can occur. This allows the system to establish a baseline state and transformation history that can be used later for audit purposes without requiring the original large-volume data to be preserved intact.
Solution Approach 2:
The system creates compressed copies of the original sensor data that retain sufficient information for audit purposes. These compressed representations serve as substitutes for the full data sets, enabling verification of data integrity and machine learning model behavior without storing or processing the complete original data volumes.
2Loss of information
If the system stores complete sensor data sets for regulatory review, then auditability is improved, but storage requirements and processing overhead increase significantly
Solution Approach 1:
The system extracts only the essential information needed for audit purposes from the complete sensor data sets. By identifying and retaining key features, provenance metadata, and compressed representations that capture the essential behavior and transformations, the system eliminates the need to store and process entire data sets while maintaining sufficient information for regulatory review.
Solution Approach 2:
The system transforms the data from its original high-volume format into a compressed representation with different parameters. This parameter transformation reduces data volume significantly while preserving the critical information needed for auditing machine learning model behavior and sensor data integrity.
3Productivity
If machine learning models process sensor data in real-time, then productivity is improved, but the ability to recover original source data for compliance purposes is reduced
Solution Approach 1:
The system performs preliminary actions by creating and storing compressed data representations and provenance metadata concurrently with the machine learning processing pipeline. This allows real-time processing to continue at full speed while simultaneously preserving the necessary information for later audit and data recovery purposes.
Solution Approach 2:
The system introduces an intermediary component that captures and stores compressed representations of sensor data and transformation metadata. This intermediary layer acts as a bridge between the real-time processing pipeline and compliance requirements, enabling both high-speed processing and data recoverability without interfering with each other.
Data Source
AI summary
In one embodiment, a method for auditing the results of a machine learning model includes: retrieving a set of state estimates for original time series data values from a database under audit; reversing the state estimation computation for each of the state estimates to produce reconstituted time series data values for each of the state estimates; retrieving the original time series data values from the database under audit; comparing the original time series data values pairwise with the reconstituted time series data values to determine whether the original time series and reconstituted time series match; and generating a signal that the database under audit (i) has not been modified where the original time series and reconstituted time series match, and (ii) has been modified where the original time series and reconstituted time series do not match.


