Edge Query Planning With Summary Data for Distributed Databases
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The proliferation of data from IoT sensors and other sources in value chain networks overwhelms traditional centralized data collection methods, leading to complexity and inefficiencies in data transmission and automated decision-making.
Innovation Solution
A method for processing queries in a distributed database using edge devices, where queries are stored on a dynamic ledger, generating approximate responses based on summary data, and transmitting these responses, with the option to use a neural network for probability distribution modeling and query planning across edge devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If centralized data collection methods are used to gather data from IoT sensors and other sources in value chain networks, then comprehensive data availability is improved, but network overhead and complexity increase significantly
Solution Approach 1:
The patent segments the centralized data collection architecture into distributed edge computing nodes deployed across the value chain network. Each edge device independently processes and stores data locally, eliminating the need for continuous centralized data transmission while maintaining comprehensive data availability through distributed storage across multiple nodes.
Solution Approach 2:
The patent introduces a new dimensional approach by implementing hierarchical data storage with summary data at edge devices and detailed data available on-demand. This multi-level data organization creates an additional dimension of data access efficiency, reducing network overhead while preserving complete data availability when needed.
2Measurement precision
If all query data is retrieved from distributed storage locations, then data accuracy is improved, but response time deteriorates due to multiple data retrieval operations
Solution Approach 1:
The patent implements preliminary action by pre-computing and storing summary data at edge devices before queries are executed. This summary data includes aggregated statistics and key metrics that can be immediately returned for common query types, providing accurate responses without requiring real-time retrieval from all distributed data sources.
Solution Approach 2:
The patent applies partial action by returning summary data for frequently queried metrics while only retrieving complete detailed data when specifically requested. This selective data retrieval approach maintains accuracy for common operations while significantly reducing response time by avoiding unnecessary full data aggregations.
3Loss of time
If summary data is used to generate approximate responses, then response time is improved, but measurement precision deteriorates
Solution Approach 1:
The patent implements dynamics by making the data retrieval strategy adaptive based on query characteristics. The system dynamically determines whether to return summary data, detailed data, or a combination based on the specific query requirements, allowing flexibility in balancing response time and accuracy for different operational contexts.
Solution Approach 2:
The patent incorporates feedback mechanisms where the system learns from query patterns and accuracy requirements. When summary data proves insufficient for specific query types or accuracy thresholds, the system automatically retrieves additional detailed data, creating a feedback loop that continuously optimizes the balance between response time and measurement precision.
Data Source
AI summary
A computer-implemented method for optimizing a distributed database includes receiving, at an aggregator, one or more query logs comprising past queries received by the distributed database. The computer-implemented method includes determining, by the aggregator, common queries received by one or more edge devices. The computer-implemented method includes determining, by the aggregator, that at least one edge device was not able to respond to a common query received by the at least one edge device. The computer-implemented method includes causing, by the aggregator, data for responding to the common query to be transmitted to the at least one edge device.


