Batch-Stream Fusion Indexing for Real-Time Data Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies face challenges in fusing and analyzing real-time data and offline data due to inconsistent data calibers, non-unified semantics, and inadequate query performance, making it difficult for big data architects and engineers to integrate and process data effectively.
Innovation Solution
A method and device for batch-stream fusion that involves obtaining an index based on an input query statement, extracting a pre-computed index data segment, and updating the query result through pre- and re-computation using a unified model, statistical information, and computation engines to position and store data segments in storage media.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If batch processing systems are used to ensure data integrity and accuracy, then data processing accuracy is improved, but real-time processing capability deteriorates
Solution Approach 1:
The patent segments the data processing system into batch processing components (for accuracy) and stream processing components (for real-time capability). The batch-stream fusion architecture divides data processing into separate batch jobs and stream processing tasks that can operate independently and concurrently, allowing each component to optimize for its specific strength while contributing to overall system performance.
Solution Approach 2:
The patent merges batch processing and stream processing into a unified batch-stream fusion architecture. This combination allows the system to leverage both the accuracy of batch processing and the real-time capability of stream processing by integrating their respective data sources, processing logic, and output mechanisms into a cohesive system that delivers both precision and speed.
2Speed
If stream computing frameworks are used to improve real-time data processing, then processing speed is improved, but data integrity and consistency deteriorate
Solution Approach 1:
The patent applies preliminary action by pre-processing and validating data through batch processing before it enters the stream processing pipeline. This preliminary batch processing ensures data quality, consistency, and integrity are established upfront, allowing the subsequent stream processing to operate on already-validated data with higher confidence in its reliability.
Solution Approach 2:
The patent implements feedback mechanisms where stream processing results are continuously monitored and validated against batch processing benchmarks. This feedback loop ensures that real-time processing maintains data integrity by comparing stream outputs with batch-processed reference data, allowing for corrective actions when deviations are detected.
3Stability of the object's composition
If separate batch and stream processing systems are maintained, then system stability is improved, but system complexity deteriorates
Solution Approach 1:
The patent creates a universal batch-stream fusion architecture that can handle both batch and stream processing workloads through a single unified system. This multi-functional platform eliminates the need for completely separate batch and stream systems by providing a common infrastructure that supports both processing modes, thereby reducing overall system complexity while maintaining stability.
Solution Approach 2:
The patent introduces intermediary components such as unified data sources, common configuration management, and integrated monitoring systems that mediate between batch and stream processing operations. These intermediaries simplify the interaction between different processing modes and provide a standardized interface, reducing the complexity of managing multiple separate systems.
Data Source
AI summary
The present invention discloses a method and a device for processing information by batch-stream fusion, and a storage medium. The method comprises the following steps: Obtaining an index based on an input query statement; extracting a pre-computed index data segment based on the index as a query result; and extracting a re-computed index data segment to update the query result. The present invention solves the technical problem that real-time data and offline data are difficult to fuse and analyze.

