Stratified Unbalanced Trees for Indexing Large Data Sets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data structures for organizing large numbers of data items face scalability and data reliability challenges, particularly in maintaining relationships among data items and synchronizing replicated instances efficiently, due to limitations in handling increased data volumes and communication bandwidth constraints.

Innovation Solution

The implementation of stratified unbalanced trees, where index data structures are configured with hierarchical nodes and fingerprint values, allowing for efficient mapping of input values to data items and enabling relaxed synchronization models to reconcile differences among replicas, thereby improving fault tolerance and performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If the size of the data structure exceeds the available physical memory, then the data structure can store more relationships, but read or write throughput becomes limited by virtual memory swapping

Engineering Contradiction:
Improvedata structure sizeVSAvoidread or write throughput
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent divides the large data structure into multiple segments or partitions that can be independently managed and stored. This segmentation allows the system to work with smaller portions of data at a time, reducing the need for frequent virtual memory swapping and maintaining higher throughput even when the total data structure size exceeds physical memory capacity.

Inventive Principle:
Principle #1Segmentation

2Reliability

If data structures are replicated to reduce data loss likelihood, then data reliability improves, but synchronization overhead and communication bandwidth requirements increase

Engineering Contradiction:
Improvedata reliabilityVSAvoidsynchronization overhead
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts the synchronization problem from the data replication process by implementing selective synchronization mechanisms. Instead of synchronizing entire data structures, the system identifies and synchronizes only the specific segments or portions that have changed, reducing the communication overhead and computational complexity while maintaining data reliability across replicas.

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If brute-force synchronization is used to ensure replica consistency, then data consistency is maintained, but substantial time and bandwidth are required to transmit and reconcile data

Engineering Contradiction:
Improvereplica consistencyVSAvoidsynchronization time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies partial action by implementing incremental synchronization that transmits and reconciles only the necessary portions of data between replicas. Instead of performing complete brute-force synchronization of entire data structures, the system identifies changed segments and synchronizes only those portions, significantly reducing the time and bandwidth required while maintaining replica consistency.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS7702640B1Stratified unbalanced trees for indexing of data items within a computer system
Publication Date: 2010.04.20 AMAZON TECH INC
  • US7702640B1 patent drawing
  • US7702640B1 patent drawing
  • US7702640B1 patent drawing

AI summary

According to one embodiment, a system may include a number of computing nodes configured to implement a number of index data structures each configured to map ones of a plurality of input values to one or more corresponding data items. Each of the index data structures may include a respective plurality of index nodes arranged hierarchically and each having an associated tag value, where each of the data items corresponds to a respective one of the index nodes, and where for a given one of the data items having a given corresponding index node, each tag value associated with each ancestor of the given corresponding index node is a prefix of a corresponding input value mapping to the given data item.