Data Replica Selector for Resilient Network Storage

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data replication methods fail to jointly consider resiliency and communication cost, particularly in the context of catastrophic concurrent failures, leading to inefficient data availability during disaster recovery.

Innovation Solution

A computer-implemented method for selecting replication nodes in a network that determines eligible nodes based on communication costs and probabilities of concurrent failure, optimizing the placement of data replicas by factoring physical distance, electrical pathway distance, and other communication parameters to minimize costs while ensuring high availability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If data is replicated on nodes close to the data source (within the same LAN or building site), then replication cost is reduced, but resiliency to catastrophic failures is compromised

Engineering Contradiction:
Improvereplication costVSAvoidresiliency to catastrophic failures
Core Design Contradiction:
Loss of energyVSReliability

Solution Approach 1:

The patent changes the selection parameters for replication nodes by introducing a composite scoring mechanism that evaluates both communication cost metrics (distance, latency, bandwidth) and resiliency metrics (geographic diversity, infrastructure independence). This multi-parameter approach allows the system to select nodes that optimize the balance between cost and reliability, rather than prioritizing one factor alone.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If data is replicated on remote, geographically diverse sites, then resiliency to catastrophic failures is improved, but replication cost increases

Engineering Contradiction:
Improveresiliency to catastrophic failuresVSAvoidreplication cost
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system transforms the node selection process by evaluating multiple parameters simultaneously including communication cost (distance, latency, bandwidth) and resiliency factors (geographic diversity, infrastructure independence). This multi-dimensional parameter evaluation enables selection of remote nodes that provide adequate protection while minimizing unnecessary replication costs.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If data is replicated on random nodes in peer-to-peer networks, then geographic diversity is achieved, but communication cost and delay increase significantly

Engineering Contradiction:
Improvegeographic diversityVSAvoidcommunication cost
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent introduces a composite scoring system that evaluates candidate nodes based on multiple parameters including communication cost metrics (distance, latency, bandwidth) and resiliency metrics (geographic diversity, infrastructure independence). This structured multi-parameter evaluation replaces random selection, ensuring that geographically diverse nodes are chosen while minimizing communication costs and delays.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS7650529B2Data replica selector
Publication Date: 2010.01.19 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US7650529B2 patent drawing
  • US7650529B2 patent drawing
  • US7650529B2 patent drawing

AI summary

There is provided a method and system for replicating data at another location. The system includes a source node that contains data in a data storage area. The source node is coupled to a network of potential replication nodes. The processor determines at least two eligible nodes in the network of nodes and determines the communication cost associated with a each of the eligible nodes. The processor also determines a probability of a concurrent failure of the source node and each of eligible nodes, and selects at least one of the eligible nodes for replication of the data located on the source node. The selection is based on the determined communication costs and probability of concurrent failure.