Distributed Data Stream Processing via Direct Streaming and Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing distributed data stream processing systems face challenges in achieving real-time processing due to their reliance on shared memory models, which lead to delays and inefficiencies, especially when handling large volumes of data, and are limited by the framework of functional reproduction in parallel computing, making it difficult to implement precise processing and maintenance.
Innovation Solution
The proposed method divides data streams into chronological segments (real-time, within 30 days, and exceeding 30 days) for parallel processing, using a multilevel framework that performs horizontal and vertical divisions based on time and data dimensions, allowing for more precise and efficient processing without relying on shared memory models, enabling real-time data processing and system responsiveness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a shared memory model is used for data exchange between modules, then data integration can be achieved, but real-time processing capability deteriorates due to delays in reading and writing data to memory
Solution Approach 1:
The patent extracts the data exchange function from the shared memory model and replaces it with direct data streaming. Instead of writing results to memory and reading them by downstream modules, the system establishes direct data flow paths where upstream modules stream data directly to downstream modules, eliminating memory I/O delays while maintaining data integration through the directed acyclic graph structure
Solution Approach 2:
The patent introduces a data exchange network as an intermediary infrastructure that enables direct communication between computing modules without relying on shared memory. This network acts as a mediator that facilitates real-time data transmission while maintaining the logical data integration required by the system architecture
2Device complexity
If functional reproduction framework is used for parallel computing, then system simplicity is maintained, but processing precision deteriorates and modularization becomes impossible
Solution Approach 1:
The patent segments the parallel computing system into heterogeneous functional modules with different processing capabilities rather than using identical computing units. Each module can be optimized for specific processing tasks (e.g., filtering, aggregation, analysis), enabling both high processing precision and modularization while maintaining manageable system complexity through standardized interfaces
Solution Approach 2:
The patent applies local quality by allowing different computing modules to have different functions and processing characteristics tailored to their specific roles in the data stream processing pipeline. This enables each module to be optimized for its local processing requirements, improving overall processing precision while maintaining system organization through the directed acyclic graph framework
3Productivity
If all data is processed in real-time, then responsiveness to user actions is improved, but processing time for large data volumes increases
Solution Approach 1:
The patent segments the data stream by time attributes into different processing queues (real-time queue for recent data, historical queue for older data). This allows the system to prioritize processing of time-sensitive data while deferring or batching processing of less time-critical data, reducing overall processing time while maintaining responsiveness to urgent user actions
Solution Approach 2:
The patent applies partial action by processing only the most recent and time-critical data in real-time through the real-time processing queue, while historical data is processed with lower priority or in batches. This selective real-time processing maintains system responsiveness for current user actions without incurring the prohibitive cost of real-time processing of all historical data
Data Source
AI summary
Embodiments of the present application relate to a distributed data stream processing method, a distributed data stream processing device, a computer program product for processing a raw data stream and a distributed data stream processing system. A distributed data stream processing method is provided. The method includes dividing a raw data stream into a real-time data stream and historical data streams, processing the real-time data stream and the historical data streams in parallel, separately generating respective results of the processing of the real-time data stream and the historical data streams, and integrating the generated processing results.


