Downstream Selective Data Backup for Stream Processing Throughput
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current distributed stream processing systems face inefficiencies in data processing due to high storage overheads and low processing speeds, particularly in upstream backup solutions where all data is backed up, constraining processing speed and affecting fault tolerance.
Innovation Solution
A downstream backup method is introduced, where selective persistence backup is performed on data and state information based on thresholds, allowing for efficient use of processing capability and improving throughput by backing up only necessary data and state changes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If upstream backup solution is used to back up all data, then system reliability is improved, but data processing efficiency and throughput are reduced
Solution Approach 1:
The patent extracts only the necessary data and state information that needs to be backed up, rather than backing up all data. This is achieved by identifying and backing up only the minimal required subset of data and state changes, thereby reducing storage overhead and improving processing efficiency while maintaining system reliability.
Solution Approach 2:
The patent applies partial action by backing up only a portion of the data - specifically, only the necessary data and state information - rather than performing complete backup of all data. This selective backup approach reduces the burden on processing resources while ensuring sufficient reliability for fault tolerance.
2Productivity
If selective persistence backup is performed on data and state information, then backup efficiency is improved and throughput is increased, but system reliability may be compromised
Solution Approach 1:
The patent changes the parameters of the backup operation by selecting specific data and state information based on predefined criteria, rather than uniformly backing up all data. This parameter-based selection approach optimizes the backup process to include only necessary information, improving efficiency while maintaining reliability through intelligent parameter selection.
Data Source
Figure 1
Figure 2
Figure 3~5
AI summary
This application discloses a data processing method and apparatus. In the data processing method, data and state information formed in a process of performing computing processing on the data are not all backed up, and a downstream computing unit performs selective persistence backup on the data and the state information. This can improve backup efficiency and reduce a total backup amount of data and state information. In this way, a processing capability of a system can be used more efficiently, and an overall throughput rate of the system can be improved.