Hub Device Re-execution Using Lineage Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In scenarios where data from geographically distributed sources, such as IoT devices, cannot be continuously transmitted to a central location quickly enough to meet Service Level Agreements or are inefficiently processed due to re-execution of analytical processes at inappropriate locations, leading to inefficiencies and non-desirable outcomes.
Innovation Solution
The implementation of a system that uses lineage metadata to identify and re-execute analytical processes on a hub device, acquiring input data from edge devices based on stored lineage metadata, allowing for efficient data processing and reducing the need for continuous data transmission.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data is continuously transmitted from geographically distributed sources to a central location, then data availability for processing is improved, but transmission time and network bandwidth consumption increase
Solution Approach 1:
The system performs preliminary actions by caching data and lineage metadata at edge devices before central processing is needed. This allows the hub device to quickly retrieve previously cached data rather than waiting for continuous transmission from distributed sources, thereby reducing transmission time while maintaining data availability.
Solution Approach 2:
The patent implements local quality by allowing different devices in the distributed system to have different data storage capabilities. Edge devices store locally relevant data and lineage metadata, while the hub device stores aggregated lineage metadata. This localizes data access and reduces the need for continuous central transmission.
2Measurement precision
If analytical processes are re-executed at the original edge device location, then data processing accuracy is maintained, but processing efficiency and resource utilization deteriorate
Solution Approach 1:
The system creates copies of data and lineage metadata and stores them at different locations (edge devices and hub device). This allows analytical processes to be re-executed at the hub device using cached copies rather than requiring execution at the original edge device, thereby improving resource utilization and processing efficiency while maintaining accuracy through the use of identical cached data.
3Loss of information
If complete lineage metadata is stored for all analytical processes, then data traceability and process re-execution capability are improved, but storage requirements and system complexity increase
Solution Approach 1:
The patent extracts and stores only the essential lineage metadata information needed for process re-execution at the hub device, rather than storing complete detailed metadata for all analytical processes. This selective extraction maintains sufficient data traceability while reducing storage requirements and simplifying the system architecture.
Data Source
AI summary
Examples disclosed herein relate to re-execution of an analytical process based on lineage metadata. In an example, a determination may be made on a hub device that an analytical process previously executed on a remote edge device is to be re-executed on the hub device, wherein the analytical process is part of an analytical workflow that is implemented at least in part on the hub device and the remote edge device. In response to the determination, a storage location of input data for re-executing the analytical process may be identified based on lineage metadata stored on the hub device, and input data may be acquired from the storage location.


