Data Stream Management System Load Shedding via Location Utility
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data stream management systems (DSMS) face challenges in maintaining quality of service (QoS) during data overload, particularly due to unpredictable data arrival rates and bursty data streams, leading to increased latency and potential for outdated or incorrect query results, which existing load shedding methods struggle to address effectively.
Innovation Solution
A dynamic load shedding scheme that assigns utility values to data stream sources based on location information, allowing the system to autonomously learn and identify low-utility data sources during normal operation, enabling targeted data discard during overload situations without requiring manual configuration or significant user input.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If load shedding is performed to handle data overload, then productivity is improved by reducing processing backlog, but reliability deteriorates by discarding potentially important data
Solution Approach 1:
The patent applies local quality by assigning different utility values to data from different sources based on their location characteristics. Instead of uniform load shedding, the system selectively discards low-utility data while preserving high-utility data, thereby maintaining reliability while improving productivity during overload conditions
2Reliability
If manual configuration of utility functions is implemented for load shedding, then reliability is improved by allowing precise control over data selection, but device complexity increases due to configuration requirements
Solution Approach 1:
The patent implements self-service by enabling the DSMS to autonomously learn and determine utility values for different data sources based on observed data patterns and QoS requirements. The system automatically adapts to changing conditions without requiring manual configuration or user input, thereby maintaining reliability while reducing device complexity
Data Source
AI summary
A data stream management system (DSMS) receives an input data stream from data stream sources and respective location information associated with sets of the data stream sources. A continuous query is executed against data items received via the input data streams to generate at least one client output data stream. A load shedding process is executed when the DSMS is overloaded with data from the input data streams. When the DSMS is not overloaded and for the location information associated with each of the data stream source sets, a respective utility value is determined indicating a utility to the client of data from the data stream source sets. The location information is stored in association with the corresponding data utility value. The location information received when the DSMS is overloaded is used, together with the data utility values, to identify input data streams whose data items are to be discarded.


