Enterprise Backup Deduplication via Social Network Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In distributed enterprise systems, conventional data backup methods consume significant bandwidth and resources, particularly due to the challenge of monitoring file movement and deduplication in a manner that balances confidentiality and resource efficiency, especially when files need to be transferred to a central server for processing.

Innovation Solution

A method and system for generating a set of unique files based on an ordered function, identifying client systems with the maximum number of files for backup, and iteratively adding them to a backup set while redefining the set of unique files and client systems, with deduplication performed at both client and enterprise levels using social network analysis and hash techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If files are transferred to a central server for backup processing, then data backup can be performed, but bandwidth consumption increases enormously

Engineering Contradiction:
Improvedata backupVSAvoidbandwidth consumption
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent segments the centralized backup process into distributed backup operations at client systems. Each client system independently performs backup of its local files using local storage resources, eliminating the need to transfer all files to a central server. This segmentation reduces network bandwidth consumption while maintaining backup reliability through distributed execution.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension of operation by performing backup at the client level rather than centralizing it. By adding the client system as an active backup participant with local storage capabilities, the system transforms the single-point centralized backup into a multi-point distributed backup architecture, reducing dependency on central server resources and network bandwidth.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If backup is performed on the central repository, then data can be backed up, but files need to be transferred before processing which consumes enormous bandwidth

Engineering Contradiction:
Improvedata backupVSAvoidfile transfer time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by performing backup operations directly at the client system before files would need to be transferred to a central server. The backup is executed in advance using local resources, eliminating the subsequent transfer step and reducing total backup time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts the backup operation from the central server environment and relocates it to the client system. By taking out the backup function from the centralized repository and implementing it locally, the system eliminates the file transfer step entirely, reducing both bandwidth consumption and time loss.

Inventive Principle:
Principle #2Taking out (Extraction)

3Quantity of substance

If deduplication is performed in real time on the central server, then data redundancy is reduced, but the process consumes more time

Engineering Contradiction:
Improvedata redundancyVSAvoiddeduplication time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent segments the deduplication process into distributed operations performed at each client system rather than centralized real-time processing on the server. Each client independently identifies and removes duplicate files locally, reducing the time burden on the central server while still achieving redundancy reduction across the network.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements partial deduplication at the client level, focusing on local duplicate files rather than attempting comprehensive real-time deduplication across all files in the network. This partial action approach reduces time consumption while still achieving significant redundancy reduction.

Inventive Principle:
Principle #16Partial or excessive action

4Loss of information

If monitoring is performed on file movement in the enterprise, then file tracking is enabled, but resource consumption increases

Engineering Contradiction:
Improvefile trackingVSAvoidresource consumption
Core Design Contradiction:
Loss of informationVSLoss of energy

Solution Approach 1:

The patent implements self-service monitoring where client systems automatically track and report their own file movements and backup status without requiring intensive centralized monitoring resources. Each client maintains local records of file operations and communicates only essential information to the server, reducing overall resource consumption while maintaining tracking capability.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10198322B2Method and system for efficient selective backup strategy in an enterprise
Publication Date: 2019.02.05 TATA CONSULTANCY SERVICES LTD
  • US10198322B2 patent drawing
  • US10198322B2 patent drawing
  • US10198322B2 patent drawing

AI summary

The present disclosure provides systems and methods to optimize data backup in a distributed enterprise system by firstly generating a set of unique files from all the files available in the enterprise. A backup set comprising files to be backed up are then generated from the set of unique files and backup is scheduled in the order in which the files to be backed up are identified. Unique files are generated based on file sharing patterns and communications among users that enable generating a social network graph from which one or more communities can be detected and deduplication can be performed on the files hosted by client systems in these communities thereby conserving resources.