Deduplicated Data Processing Congestion Control via PID Setpoints
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing deduplication systems face challenges in controlling congestion to prevent failures and timeouts during data processing, while maximizing system resource utilization, especially in dynamic environments where resource access patterns are unknown.
Innovation Solution
A method for deduplicated data processing congestion control using a congestion target setpoint calculated with proportional, integral, and derivative constants, which adjusts the number of active deduplicated data processes to maintain optimal resource utilization without prior knowledge of resource access patterns, ensuring continuous operation without failures or timeouts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If deduplicated data processing activity is increased to maximize productivity, then system resource utilization is improved, but congestion increases causing failures or timeouts
Solution Approach 1:
The patent implements dynamic congestion control by continuously adjusting the degree of deduplication data processing activity based on real-time congestion metrics. The system dynamically modifies process parameters and resource allocation to adapt to changing congestion conditions, allowing the system to operate at optimal productivity levels while preventing failures through continuous adjustment rather than static operation.
Solution Approach 2:
The patent employs feedback mechanisms where congestion metrics are monitored and fed back into the control system. This feedback loop enables the system to detect congestion conditions and automatically adjust processing activity to maintain reliability. The feedback from congestion monitoring allows the system to modulate data processing intensity to prevent timeouts and failures while maximizing resource utilization.
2Reliability
If conventional congestion control approaches are used, then some congestion management is achieved, but system resource utilization is not maximized
Solution Approach 1:
The patent changes key parameters of the deduplication processing system dynamically based on congestion conditions. By modifying processing parameters such as data block sizes, indexing intensity, and resource allocation ratios, the system optimizes the balance between congestion control and resource utilization. These parameter changes enable the system to maintain high productivity while achieving effective congestion management.
3Measurement precision
If prior knowledge of resource access patterns is required for congestion control, then control accuracy is improved, but system adaptability to dynamic environments decreases
Solution Approach 1:
The patent implements self-service congestion control where the system automatically monitors its own congestion state and adjusts processing activity without requiring external input or prior knowledge of resource access patterns. The system serves itself by using internal congestion metrics to trigger adaptive control actions, enabling it to adapt to dynamic environments autonomously while maintaining accurate congestion control through real-time self-monitoring.
Data Source
AI summary
Various embodiments for deduplicated data processing congestion control in a computing environment are provided. In one such embodiment, a congestion target setpoint is calculated using one of a proportional constant, an integral constant, and a derivative constant, wherein the congestion target setpoint is a virtual dimension setpoint. A single congestion metric is determined from a sampling of a plurality of combined deduplicated data processing congestion statistics in a number of active deduplicated data processes. A congestion limit is calculated from a comparison of the single congestion metric to the congestion target setpoint, the congestion limit being a manipulated variable. The congestion limit is compared to the number of active deduplicated data processes. If the number of active deduplicated data processes are less than the congestion limit, a new deduplicated data process of the number of active deduplicated data processes is spawned.


