Distributed Data Stream Processing via Direct Streaming and Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing distributed data stream processing systems face challenges in achieving real-time processing due to their reliance on shared memory models, which lead to delays and inefficiencies, especially when handling large volumes of data, and are limited by the framework of functional reproduction in parallel computing, making it difficult to implement precise processing and maintenance.

Innovation Solution

The proposed method divides data streams into chronological segments (real-time, within 30 days, and exceeding 30 days) for parallel processing, using a multilevel framework that performs horizontal and vertical divisions based on time and data dimensions, allowing for more precise and efficient processing without relying on shared memory models, enabling real-time data processing and system responsiveness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a shared memory model is used for data exchange between modules, then data integration can be achieved, but real-time processing capability deteriorates due to delays in reading and writing data to memory

Engineering Contradiction:
Improvedata integrationVSAvoidreal-time processing speed
Core Design Contradiction:
ReliabilityVSSpeed

Solution Approach 1:

The patent extracts the data exchange function from the shared memory model and replaces it with direct data streaming. Instead of writing results to memory and reading them by downstream modules, the system establishes direct data flow paths where upstream modules stream data directly to downstream modules, eliminating memory I/O delays while maintaining data integration through the directed acyclic graph structure

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces a data exchange network as an intermediary infrastructure that enables direct communication between computing modules without relying on shared memory. This network acts as a mediator that facilitates real-time data transmission while maintaining the logical data integration required by the system architecture

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If functional reproduction framework is used for parallel computing, then system simplicity is maintained, but processing precision deteriorates and modularization becomes impossible

Engineering Contradiction:
Improvesystem framework simplicityVSAvoidprocessing precision
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent segments the parallel computing system into heterogeneous functional modules with different processing capabilities rather than using identical computing units. Each module can be optimized for specific processing tasks (e.g., filtering, aggregation, analysis), enabling both high processing precision and modularization while maintaining manageable system complexity through standardized interfaces

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by allowing different computing modules to have different functions and processing characteristics tailored to their specific roles in the data stream processing pipeline. This enables each module to be optimized for its local processing requirements, improving overall processing precision while maintaining system organization through the directed acyclic graph framework

Inventive Principle:
Principle #3Local quality

3Productivity

If all data is processed in real-time, then responsiveness to user actions is improved, but processing time for large data volumes increases

Engineering Contradiction:
Improvesystem responsivenessVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent segments the data stream by time attributes into different processing queues (real-time queue for recent data, historical queue for older data). This allows the system to prioritize processing of time-sensitive data while deferring or batching processing of less time-critical data, reducing overall processing time while maintaining responsiveness to urgent user actions

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by processing only the most recent and time-critical data in real-time through the real-time processing queue, while historical data is processed with lower priority or in batches. This selective real-time processing maintains system responsiveness for current user actions without incurring the prohibitive cost of real-time processing of all historical data

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9727613B2Distributed data stream processing method and system
Publication Date: 2017.08.08 ALIBABA GROUP HOLDING LTD
  • US9727613B2 patent drawing
  • US9727613B2 patent drawing
  • US9727613B2 patent drawing

AI summary

Embodiments of the present application relate to a distributed data stream processing method, a distributed data stream processing device, a computer program product for processing a raw data stream and a distributed data stream processing system. A distributed data stream processing method is provided. The method includes dividing a raw data stream into a real-time data stream and historical data streams, processing the real-time data stream and the historical data streams in parallel, separately generating respective results of the processing of the real-time data stream and the historical data streams, and integrating the generated processing results.