Multidimensional Data Replica Selector Optimizing Resiliency and Cost
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data replication solutions fail to jointly consider resiliency and replication costs, leading to inefficient data availability in failure recovery scenarios, as they either prioritize geographic proximity or diversity without accounting for correlated failures and communication costs.
Innovation Solution
A computer-implemented method that constructs a multidimensional model to select replication nodes based on system characteristics such as geographic location, administrative domain, hardware, and network type, determining data availability and replication costs to optimize node selection for data replication, balancing resiliency and cost effectively.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If data is replicated on geographically close nodes, then communication replication cost is reduced, but geographic diversity is insufficient to survive catastrophic failures
Solution Approach 1:
The patent changes the selection parameters for replication nodes from simple geographic proximity to a multidimensional model that includes geographic location, administrative domain, hardware type, operating system, and network type. This allows the system to select nodes that optimize both cost and resiliency by considering multiple dimensions simultaneously rather than relying solely on geographic distance.
Solution Approach 2:
The patent introduces additional dimensions beyond geographic location, including administrative domain, hardware characteristics, software environment, and network infrastructure. By evaluating nodes across multiple dimensions, the system can identify replicas that are cost-effective while simultaneously providing diverse protection against various failure modes including catastrophic events.
2Reliability
If data is replicated on remote geographically diverse sites, then resiliency to catastrophes is improved, but communication cost and infrastructure cost increase
Solution Approach 1:
The patent modifies the cost and availability calculation to incorporate multiple dimensions including geographic location, administrative domain, hardware type, operating system, and network type. This enables the system to identify intermediate solutions that may not be the most geographically distant but provide sufficient diversity across multiple dimensions while reducing communication costs.
Solution Approach 2:
The patent applies different selection criteria to different dimensions of node characteristics. Rather than uniformly prioritizing geographic distance, the system evaluates each dimension (geographic location, administrative domain, hardware, software, network) and selects nodes that provide optimal local quality in each dimension, resulting in an overall optimized replication strategy.
3Reliability
If peer-to-peer systems replicate content across multiple nodes, then data availability is improved, but communication delays and costs increase due to random node selection
Solution Approach 1:
The patent incorporates feedback mechanisms that evaluate node characteristics across multiple dimensions and use this information to make informed replication decisions. The system continuously assesses the multidimensional model and adjusts replication strategies based on the evaluated data availability and cost metrics, avoiding random selection and reducing communication delays.
Solution Approach 2:
The patent performs preliminary evaluation of candidate nodes across multiple dimensions before selecting replication targets. By pre-assessing geographic location, administrative domain, hardware compatibility, software environment, and network characteristics, the system identifies optimal replication nodes in advance, avoiding the need for time-consuming random selection and communication delays during failure events.
Data Source
AI summary
A method is provided for selecting a replication node from eligible nodes in a network. A multidimensional model is constructed that defines a multidimensional space and includes the eligible nodes, with each of the dimensions of the multidimensional model being a system characteristic. A data availability value is determined for each of the eligible nodes, and a cost of deploying is determined for each of at least two availability strategies to the eligible nodes. At least one of the eligible nodes is selected for replication of data that is stored on a source node in the network. The selecting step includes selecting the eligible node whose: data availability value is determined to be highest among the eligible nodes whose cost of deploying does not exceed a specified maximum, or cost of deploying is determined to be lowest among the eligible nodes whose data availability value does not exceed a specified minimum.


