Backup Data Deduplication Scheduling Under Compute Resource Constraints
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data backup methods, relying on data deduplication and compression, consume significant computing resources and are limited by the availability of computing resources, resulting in low backup data amounts and inefficient use of time resources, especially during off-peak hours.
Innovation Solution
A data processing method that dynamically determines deduplication and compression based on computing resource utilization and deduplication/compression rates, optimizing these processes to reduce resource overhead, increase backup data amount, and ensure continuous production service by performing them during low resource utilization periods.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Volume of stationary object
If data deduplication and compression are performed during backup, then storage space is reduced, but computing resource consumption increases
Solution Approach 1:
The patent implements dynamic adjustment of deduplication and compression operations based on real-time computing resource availability. The system monitors computing resource usage and dynamically decides whether to perform deduplication, compression, or both operations during backup, transforming a static resource consumption model into a dynamic one that adapts to system conditions.
Solution Approach 2:
The system changes operational parameters (deduplication rate, compression rate) based on computing resource status. When computing resources are abundant, higher deduplication and compression rates are applied; when resources are constrained, the system reduces or skips these operations, thereby adjusting the balance between storage efficiency and resource consumption.
2Volume of stationary object
If data deduplication and compression are performed during backup, then storage space is reduced, but backup data amount decreases
Solution Approach 1:
The system dynamically adjusts the level of deduplication and compression based on computing resource availability. When resources are abundant, the system performs both operations to maximize storage efficiency; when resources are limited, it reduces or skips these operations to preserve backup data amount, thus dynamically balancing storage efficiency against backup completeness.
3Use of energy by moving object
If backup is performed during off-peak hours, then computing resource availability increases, but time resources for backup are limited
Solution Approach 1:
The patent enables continuous backup operations by performing deduplication and compression dynamically during backup based on computing resource availability. This eliminates the need to restrict backup to specific off-peak time windows, allowing backup to proceed continuously while adapting resource usage to system conditions, thereby converting a time-constrained batch process into a continuous adaptive process.
4Volume of stationary object
If deduplication rate threshold is set high, then storage space is reduced more, but computing resource consumption increases
Solution Approach 1:
The system implements feedback control by monitoring computing resource usage and adjusting deduplication operations accordingly. The deduplication rate is not fixed but is dynamically adjusted based on feedback from system resource monitoring, allowing the system to achieve high deduplication rates when resources are abundant while reducing deduplication intensity when resources are constrained.
Data Source
AI summary
A data processing method includes a computing device that obtains first data, and determines a deduplication rate of the first data and a working status of the computing device. When the working status of the computing device is a busy state, and when the deduplication rate of the first data is greater than or equal to a first preset value, deduplication on the first data is performed. When the deduplication rate of the first data is lower than the first preset value, deduplication is not performed on the first data, and the first data may be directly stored.


