Mirrored Index Tree for Distributed Data Restoration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing distributed data management systems face challenges in effectively managing indexing data across multiple nodes, particularly in handling device failures, restoring offline data, and reintegrating devices that come back online.

Innovation Solution

The implementation of a mirrored balanced index tree structure with redundant copies of nodes stored across multiple devices, allowing for the traversal and restoration of nodes on inaccessible devices, and merging updated nodes upon reintegration.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data is stored in a distributed manner across multiple nodes, then system flexibility and reliability are improved, but complexity of managing data consistency and handling device failures increases

Engineering Contradiction:
Improvesystem reliabilityVSAvoiddata management complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The index tree is segmented into multiple nodes that are distributed across different devices. Each node contains a portion of the indexing data, and the tree structure is divided such that root nodes, internal nodes, and leaf nodes can be independently stored and managed on different devices, enabling parallel access and improved reliability

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Redundant copies of index tree nodes are created and stored on multiple devices before any failure occurs. This preliminary duplication ensures that when a device fails, the system can immediately access backup copies without needing to reconstruct the missing data, thereby maintaining reliability while managing complexity through pre-established redundancy

Inventive Principle:
Principle #10Preliminary action

2Reliability

If multiple copies of index nodes are stored across devices, then data availability is improved, but storage space consumption increases

Engineering Contradiction:
Improvedata availabilityVSAvoidstorage space
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

Different levels of the index tree are stored with different redundancy patterns based on their local quality requirements. Root nodes and internal nodes that are accessed more frequently for navigation purposes are replicated across more devices, while leaf nodes containing actual data may have different replication strategies, optimizing the balance between availability and storage consumption

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

Selective copying of index tree nodes is performed based on their importance and access patterns. Rather than copying entire trees or all nodes uniformly, the system identifies critical nodes that need redundancy for availability and copies only those specific nodes to appropriate backup devices, reducing overall storage overhead while maintaining data availability

Inventive Principle:
Principle #26Copying

3Reliability

If the index tree is traversed to locate and restore nodes on inaccessible devices, then data integrity is improved, but processing time increases

Engineering Contradiction:
Improvedata integrityVSAvoidrestoration time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system maintains metadata that preliminarily records the location and status of index tree nodes across devices. When a device becomes inaccessible, the restoration process can quickly query this pre-established metadata to identify which nodes need restoration and from which backup devices, avoiding the need to traverse the entire index tree structure and significantly reducing restoration time while maintaining data integrity

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The restoration process extracts only the specific missing nodes from the distributed system rather than attempting to restore or verify the entire index tree. By identifying and extracting only the nodes that are missing or corrupted based on metadata information, the system minimizes processing time while ensuring data integrity through targeted restoration operations

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS7797283B2Systems and methods for maintaining distributed data
Publication Date: 2010.09.14 EMC IP HLDG CO LLC
  • US7797283B2 patent drawing
  • US7797283B2 patent drawing
  • US7797283B2 patent drawing

AI summary

Systems and methods are disclosed that provide an indexing data structure. In one embodiment, the indexing data structure is mirrored index tree where the copies of the nodes of the tree are stored across devices in a distributed system. In one embodiment, nodes that are stored on an offline device are restored, and an offline device that comes back online is merged into the distributed system and given access to the current indexing data structure. In one embodiment, the indexing data structure is traversed to locate and restore nodes that are stored on offline devices of the distributed system.