Distributed Data Backup Location Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current backup technologies for distributed data often lead to increased network congestion and latency, particularly when backing up heavily utilized nodes or servers, and fail to consider the active or passive status of data sources, resulting in performance impacts and inefficient resource utilization.

Innovation Solution

A method for identifying and configuring backup storage locations based on specified preferences, including network distance, resource availability, and ability to support parallel processing, to minimize resource impact and optimize backup efficiency, which involves calculating network distance and resource utilization to select suitable backup locations for distributed data sources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If backup jobs are scheduled for heavily utilized nodes or remote locations, then backup coverage is improved, but network congestion and latency increase

Engineering Contradiction:
Improvebackup coverageVSAvoidnetwork latency
Core Design Contradiction:
ReliabilityVSSpeed

Solution Approach 1:

The patent applies local quality by selecting backup storage locations based on their proximity to data sources and current resource availability. Different backup locations are chosen for different data sources depending on local network conditions, storage capacity, and workload characteristics, rather than using a single remote backup location for all data.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically selects backup storage locations based on real-time or near-real-time conditions such as network congestion levels, storage availability, and node utilization. The backup location selection is not static but adapts to changing system conditions, allowing the system to optimize performance while maintaining backup coverage.

Inventive Principle:
Principle #15Dynamics

2Ease of operation

If backup jobs are scheduled without considering active or passive node status, then backup scheduling simplicity is maintained, but performance impact on active nodes increases

Engineering Contradiction:
Improvescheduling simplicityVSAvoidnode performance
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent incorporates feedback by monitoring the operational status of nodes (active or passive) and using this information to inform backup scheduling decisions. The system gathers information about node roles and performance characteristics, then uses this feedback to schedule backups in a way that minimizes impact on active nodes while maintaining scheduling automation.

Inventive Principle:
Principle #23Feedback

3Ease of operation

If backup storage locations are not selected based on resource availability, then backup location selection simplicity is maintained, but resource utilization efficiency decreases

Engineering Contradiction:
Improvelocation selection simplicityVSAvoidresource utilization efficiency
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The system changes the parameters used for backup location selection from simple static criteria to dynamic parameters including available storage capacity, current resource utilization, network bandwidth availability, and node performance metrics. These parameter changes enable the system to optimize resource utilization while maintaining automated selection processes.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8140791B1Techniques for backing up distributed data
Publication Date: 2012.03.20 COHESITY INC
  • US8140791B1 patent drawing
  • US8140791B1 patent drawing
  • US8140791B1 patent drawing

AI summary

Techniques for backing up distributed data are disclosed. In one particular exemplary embodiment, the techniques may be realized as a method for backing up distributed data comprising identifying one or more sources of distributed data targeted for backup, identifying two or more backup storage locations, determining which one or more backup storage locations of the two or more identified backup storage locations to utilize for a backup job based at least in part on one or more specified preferences, and configuring, for at least one of the sources of distributed data, the backup job using the one or more backup storage locations.