Enterprise Backup Deduplication via Social Network Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In distributed enterprise systems, conventional data backup methods consume significant bandwidth and resources, particularly due to the challenge of monitoring file movement and deduplication in a manner that balances confidentiality and resource efficiency, especially when files need to be transferred to a central server for processing.
Innovation Solution
A method and system for generating a set of unique files based on an ordered function, identifying client systems with the maximum number of files for backup, and iteratively adding them to a backup set while redefining the set of unique files and client systems, with deduplication performed at both client and enterprise levels using social network analysis and hash techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If files are transferred to a central server for backup processing, then data backup can be performed, but bandwidth consumption increases enormously
Solution Approach 1:
The patent segments the centralized backup process into distributed backup operations at client systems. Each client system independently performs backup of its local files using local storage resources, eliminating the need to transfer all files to a central server. This segmentation reduces network bandwidth consumption while maintaining backup reliability through distributed execution.
Solution Approach 2:
The patent introduces a new dimension of operation by performing backup at the client level rather than centralizing it. By adding the client system as an active backup participant with local storage capabilities, the system transforms the single-point centralized backup into a multi-point distributed backup architecture, reducing dependency on central server resources and network bandwidth.
2Reliability
If backup is performed on the central repository, then data can be backed up, but files need to be transferred before processing which consumes enormous bandwidth
Solution Approach 1:
The patent applies preliminary action by performing backup operations directly at the client system before files would need to be transferred to a central server. The backup is executed in advance using local resources, eliminating the subsequent transfer step and reducing total backup time.
Solution Approach 2:
The patent extracts the backup operation from the central server environment and relocates it to the client system. By taking out the backup function from the centralized repository and implementing it locally, the system eliminates the file transfer step entirely, reducing both bandwidth consumption and time loss.
3Quantity of substance
If deduplication is performed in real time on the central server, then data redundancy is reduced, but the process consumes more time
Solution Approach 1:
The patent segments the deduplication process into distributed operations performed at each client system rather than centralized real-time processing on the server. Each client independently identifies and removes duplicate files locally, reducing the time burden on the central server while still achieving redundancy reduction across the network.
Solution Approach 2:
The patent implements partial deduplication at the client level, focusing on local duplicate files rather than attempting comprehensive real-time deduplication across all files in the network. This partial action approach reduces time consumption while still achieving significant redundancy reduction.
4Loss of information
If monitoring is performed on file movement in the enterprise, then file tracking is enabled, but resource consumption increases
Solution Approach 1:
The patent implements self-service monitoring where client systems automatically track and report their own file movements and backup status without requiring intensive centralized monitoring resources. Each client maintains local records of file operations and communicates only essential information to the server, reducing overall resource consumption while maintaining tracking capability.
Data Source
AI summary
The present disclosure provides systems and methods to optimize data backup in a distributed enterprise system by firstly generating a set of unique files from all the files available in the enterprise. A backup set comprising files to be backed up are then generated from the set of unique files and backup is scheduled in the order in which the files to be backed up are identified. Unique files are generated based on file sharing patterns and communications among users that enable generating a social network graph from which one or more communities can be detected and deduplication can be performed on the files hosted by client systems in these communities thereby conserving resources.


