Management Server Data Migration Duplicate Removal
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In large-scale systems with multiple storage apparatuses, the addition of new storage apparatuses and backup servers to manage backup processing load results in inefficient data duplicate removal, leading to increased backup data volume and prolonged processing times, especially when data duplication rates between storage apparatuses are low.
Innovation Solution
A management server is implemented to calculate backup capacity and determine migration targets and destinations based on data duplication rates, allowing data with low duplication rates to be migrated to storage apparatuses with higher duplication rates, thereby improving duplicate removal efficiency and reducing backup processing time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data duplicate removal technology is applied in each storage apparatus, then backup processing load on each backup server is reduced, but backup data volume increases and backup processing time exceeds expected time limit when data duplication rate is low
Solution Approach 1:
The patent merges duplicate removal processing across multiple storage apparatuses by introducing a management server that coordinates between storage apparatuses. Instead of each storage apparatus performing duplicate removal independently (which fails when duplication rates are low), the system combines data from multiple storage apparatuses to find duplicates across the entire system, thereby maintaining low backup data volume while distributing processing load.
Solution Approach 2:
The management server acts as an intermediary between storage apparatuses and backup servers. It receives data from multiple storage apparatuses, performs coordinated duplicate removal processing, and manages the backup process. This intermediary role enables the system to achieve both load distribution and effective duplicate removal by centralizing the duplicate detection logic while distributing data sources.
2Productivity
If backup servers are added to distribute backup processing load, then backup processing capacity increases, but system complexity and cost increase
Solution Approach 1:
The management server performs multiple functions: it manages storage apparatuses, coordinates duplicate removal processing across multiple storage systems, and controls the backup process. This multi-functional approach eliminates the need for separate backup servers, achieving load distribution and increased backup capacity without adding dedicated backup infrastructure, thereby controlling system complexity.
3Quantity of substance
If data with low duplication rate is backed up, then backup data volume increases, but duplicate removal efficiency decreases
Solution Approach 1:
The patent extends the duplicate removal search from a single storage apparatus dimension to a multi-storage-apparatus dimension. By searching across multiple storage apparatuses simultaneously, the system finds duplicates that would be missed in individual apparatuses, even when local duplication rates are low. This dimensional expansion maintains high duplicate removal efficiency while managing backup data volume effectively.
Data Source
AI summary
A management server and a data migration method shorten a data backup processing time and improve the efficiency of data duplicate removal in a storage apparatus. The management server includes a backup capacity calculation unit for calculating, based on duplicate information of data stored in a plurality of storage apparatuses, a backup capacity of the data if duplicate data is removed. The management server includes a migration source data determination unit for determining, based on the calculated backup capacity, data which is a target for migration to another storage apparatus among data stored in one of the storage apparatuses. The management server further includes a migration destination storage apparatus determination unit for determining, based on duplicate information of the migration target data and data of a storage apparatus which is a migration destination for the data, a migration destination storage apparatus for the data.


