Data Discovery System Backpressure Control for Latency Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In data discovery processes, the mismatch in speeds between data collection and processing tools often results in missed data and latency, as collection tools operate faster than processing tools, leading to data drop and user latency.
Innovation Solution
A system that includes a crawler program, a data fetcher program, and circuit hardware to regulate speeds through 'backpressure' mechanisms, ensuring that data is processed efficiently by coordinating the speeds of different data discovery tasks, preventing data drop and latency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If the identification-collection tool collects data at high speed, then data collection efficiency is improved, but the processing tool cannot keep up, resulting in data loss and latency
Solution Approach 1:
The system implements feedback mechanisms where the processing tool communicates its processing capacity and status back to the identification-collection tool. This allows the collection tool to adjust its data collection rate dynamically based on the processing tool's current load and throughput, ensuring that data is collected at a sustainable rate that prevents overflow and loss while maximizing collection efficiency.
Solution Approach 2:
The system transitions from static, fixed-speed data collection to dynamic, adaptive data collection. The identification-collection tool continuously adjusts its operating parameters based on real-time feedback from the processing tool, allowing the system to optimize data collection speed while maintaining reliability under varying workload conditions.
2Speed
If the identification-collection tool collects data faster than the processing tool can process, then data collection speed is improved, but latency increases and critical data may be dropped
Solution Approach 1:
The feedback loop enables the identification-collection tool to receive real-time information about processing queue depth and current processing speed. Based on this feedback, the system dynamically adjusts the data collection rate to match processing capacity, preventing the accumulation of unprocessed data that would increase latency and ensure critical data is processed in a timely manner.
Solution Approach 2:
The system enables self-regulation where the data collection process automatically adjusts its own speed based on system conditions. The identification-collection tool monitors the overall system state and autonomously modulates its data collection rate without requiring external intervention, optimizing the balance between collection speed and processing latency.
3Ease of operation
If manual transfer is used between data discovery tools, then data transfer is performed, but substantial errors occur in the tools and data discovery process
Solution Approach 1:
The system merges the identification-collection tool and processing tool into an integrated data discovery system with standardized internal interfaces. This unification eliminates the need for manual data transfer between separate tools, creating a seamless data flow that maintains data integrity and reduces errors while preserving operational simplicity through a unified user interface.
Solution Approach 2:
The system introduces standardized data interfaces and protocols as intermediaries between different data discovery tools. These intermediaries provide structured, validated data exchange mechanisms that replace error-prone manual transfers, ensuring data accuracy and consistency while maintaining ease of operation through automated interface management.
Data Source
AI summary
A system for facilitating data discovery on a network, wherein the network has one or more data storage devices. The system may include a crawler program configured to select at least a first set of files and a second set of files, each of the first set of files and the second set of files being stored in at least one of the one or more data storage devices. The system may also include a data fetcher program configured to obtain a copy of the first set of files, the data fetcher program being further configured to resist against obtaining a copy of the second set of files. The system may also include circuit hardware implementing one or more functions of one or more of the crawler program and the data fetcher program.


